SiliconFlow

SiliconFlow

One-stop LLM API platform — high-speed inference, pay-as-you-go, first stop for new flagship models like GLM-5.2 and DeepSeek-V4

Visit official site →This page contains affiliate links — we may earn a commission if you sign up through them.

Why we recommend SiliconFlow

Why SiliconFlow

SiliconFlow (siliconflow.cn) is a one-stop LLM API platform: hundreds of open-source and flagship models (DeepSeek, GLM, Qwen, Kimi, MiniMax, LongCat, and more) behind a single OpenAI-compatible interface, billed pay-as-you-go, covering language, speech, image, and video. For developers who don't want to commit to one vendor — or who want to try several models with the lowest friction — it's a very cost-effective starting point.

One platform, two entry points. siliconflow.cn is the brand site — model intro, price center, and the full product matrix; the "Visit official site" button at the top points here. cloud.siliconflow.cn is the developer console — API keys, balance top-up, billing details, and the referral program; the "View plan" buttons on the right jump to the console. Same account for both: sign up once and you can use both. There's also an international site at www.siliconflow.com.

Flagship models & pricing

The pricing page (siliconflow.cn/pricing) lists per-model input / output / cache-hit rates (per 1M tokens), billed against your prepaid balance — no fixed monthly fee:

  • DeepSeek-V4-Pro — ¥12 in / ¥24 out. 1.6T MoE flagship with native 1M-token context; top-tier on coding and agentic tasks
  • GLM-5.2 (high-speed) — ¥8 in / ¥28 out. #2 on Code Arena, built for long-horizon, high-complexity software engineering
  • Kimi-K2.7-Code — ¥6.5 in / ¥27 out. A screen-reading code expert that uses ~30% fewer thinking tokens than the previous generation

As an aggregation platform, SiliconFlow is also often the worldwide first stop for high-speed versions of new models: DeepSeek-V4-Pro/Flash, GLM-5.2, Kimi-K2.7-Code, and Kimi-K3 all launched here with fast inference (official numbers: 10x+ language-model speedup, 1-second image generation, up to 66% cost savings).

Pay-as-you-go vs Coding Plan: which one?

DimensionSiliconFlow (pay-as-you-go)Typical Coding Plan (Zhipu / Qwen)
BillingPrepaid balance, per-tokenFixed monthly / annual fee
ModelsHundreds behind one APIVendor's own models
Best forMulti-model comparison, spiky usage, agent workflowsHeavy fixed subscription to one vendor
RiskWatch per-model unit priceFixed fee even when idle

If you want to compare models side by side or run an agent tool (Claude Code / Cursor / OpenCode) against multiple backends, SiliconFlow is the more flexible choice; if your usage is steady and you're committed to one vendor, a Coding Plan's monthly rate may be cheaper.

Getting started

  1. Sign up via the invite link cloud.siliconflow.cn/i/X2aUwDBp (or directly at siliconflow.cn)
  2. Top up your balance in the console (multiple payment methods, invoicing available)
  3. Create an API key
  4. Point your OpenAI SDK base URL at the SiliconFlow endpoint and pass the key
  5. First successful request → balance is deducted per token

Check the official page for live prices — per-model rates change as new models launch and promotions rotate; this page doesn't hardcode every figure. See siliconflow.cn/pricing for the latest.