SiliconFlow
One-stop LLM API platform — high-speed inference, pay-as-you-go, first stop for new flagship models like GLM-5.2 and DeepSeek-V4
Why we recommend SiliconFlow
Why SiliconFlow
SiliconFlow (siliconflow.cn) is a one-stop LLM API platform: hundreds of open-source and flagship models (DeepSeek, GLM, Qwen, Kimi, MiniMax, LongCat, and more) behind a single OpenAI-compatible interface, billed pay-as-you-go, covering language, speech, image, and video. For developers who don't want to commit to one vendor — or who want to try several models with the lowest friction — it's a very cost-effective starting point.
One platform, two entry points. siliconflow.cn is the brand site — model intro, price center, and the full product matrix; the "Visit official site" button at the top points here. cloud.siliconflow.cn is the developer console — API keys, balance top-up, billing details, and the referral program; the "View plan" buttons on the right jump to the console. Same account for both: sign up once and you can use both. There's also an international site at www.siliconflow.com.
Flagship models & pricing
The pricing page (siliconflow.cn/pricing) lists per-model input / output / cache-hit rates (per 1M tokens), billed against your prepaid balance — no fixed monthly fee:
- DeepSeek-V4-Pro — ¥12 in / ¥24 out. 1.6T MoE flagship with native 1M-token context; top-tier on coding and agentic tasks
- GLM-5.2 (high-speed) — ¥8 in / ¥28 out. #2 on Code Arena, built for long-horizon, high-complexity software engineering
- Kimi-K2.7-Code — ¥6.5 in / ¥27 out. A screen-reading code expert that uses ~30% fewer thinking tokens than the previous generation
As an aggregation platform, SiliconFlow is also often the worldwide first stop for high-speed versions of new models: DeepSeek-V4-Pro/Flash, GLM-5.2, Kimi-K2.7-Code, and Kimi-K3 all launched here with fast inference (official numbers: 10x+ language-model speedup, 1-second image generation, up to 66% cost savings).
Pay-as-you-go vs Coding Plan: which one?
| Dimension | SiliconFlow (pay-as-you-go) | Typical Coding Plan (Zhipu / Qwen) |
|---|---|---|
| Billing | Prepaid balance, per-token | Fixed monthly / annual fee |
| Models | Hundreds behind one API | Vendor's own models |
| Best for | Multi-model comparison, spiky usage, agent workflows | Heavy fixed subscription to one vendor |
| Risk | Watch per-model unit price | Fixed fee even when idle |
If you want to compare models side by side or run an agent tool (Claude Code / Cursor / OpenCode) against multiple backends, SiliconFlow is the more flexible choice; if your usage is steady and you're committed to one vendor, a Coding Plan's monthly rate may be cheaper.
Getting started
- Sign up via the invite link cloud.siliconflow.cn/i/X2aUwDBp (or directly at siliconflow.cn)
- Top up your balance in the console (multiple payment methods, invoicing available)
- Create an API key
- Point your OpenAI SDK base URL at the SiliconFlow endpoint and pass the key
- First successful request → balance is deducted per token
Check the official page for live prices — per-model rates change as new models launch and promotions rotate; this page doesn't hardcode every figure. See siliconflow.cn/pricing for the latest.