Qwen3-Coder API: 3 Ways to Use It Free in 2026

Quick answer: The old "just log in with Qwen OAuth for free" route is gone — Qwen's docs confirm that free tier was discontinued on 2026-04-15. But Qwen3-Coder is open-weights under Apache 2.0, so you can still use it at $0 two ways: self-hosting the 30B Flash model locally, or Puter.js in the browser. OpenRouter's qwen/qwen3-coder:free endpoint was retired — Qwen3-Coder is paid-only there now. Alibaba's official Model Studio API is pay-as-you-go for scale.

Qwen3-Coder is Alibaba's open-weights coding model family, released July 2025 and tuned for agentic work — reading a repo, planning, calling tools, iterating. Because the weights are open, closing one free front door didn't kill free access. Here's the honest map of what's actually free in 2026, with runnable code.

What Is Qwen3-Coder?

Three things make it stand out among open coding models:

  • MoE at frontier scale. The flagship Qwen3-Coder-480B-A35B-Instruct has 480B total parameters but only ~35B active per token (8 of 160 experts fire) — giant-model capacity at mid-sized inference cost.
  • Codebase-scale context. 256K tokens natively, up to 1M with extrapolation — enough to hold a mid-sized repo in one prompt.
  • Genuinely open weights. Apache 2.0 on Hugging Face — inspect, fine-tune, run offline, deploy commercially. That openness is what lets third parties offer free endpoints.

There's also Qwen3-Coder-Flash (Qwen3-Coder-30B-A3B-Instruct — 30B total, ~3B active), the variant most developers run locally: it fits on a single high-VRAM consumer GPU and speaks the same tool-calling format. "Run Qwen3-Coder for free on my own machine" almost always means Flash.

Benchmarks

BenchmarkQwen3-Coder-480B-A35BWhat it measures
SWE-bench Verified68.4% (342 of 500 resolved)Real GitHub bug-fix tasks — agentic coding
Context (native / extended)256K / 1M tokensHow much code fits in one prompt
ArchitectureMoE, 480B total / 35B activeCapacity vs inference cost

Alibaba positions that 68.4% as comparable to Claude Sonnet 4 on SWE-bench Verified — remarkable for an openly downloadable model. Knowledge cutoff is June 2025; released July 23, 2025.

The 3 Ways to Use Qwen3-Coder for Free

Way 1 — OpenRouter (the :free variant is gone)

Most guides still tell you to call qwen/qwen3-coder:free on OpenRouter. That ID no longer resolves. OpenRouter still carries Qwen3-Coder — qwen/qwen3-coder, qwen3-coder-plus, qwen3-coder-flash and others — but every one of them is paid. Checked against OpenRouter's public catalog on 2026-08-28, none of the 51 qwen/ models carries a :free suffix. For a genuinely $0 coding model on OpenRouter today, cohere/north-mini-code:free is the coding-focused entry on the free list; openrouter/free routes to whatever is free at the moment and never goes stale.

The trade-off is rate limits. Per OpenRouter, free models are shared and throttled — on the order of ~20 requests/minute and roughly 50 requests/day, with the daily ceiling rising to about 1,000/day once you've ever purchased $10 in credits (a one-time lifetime threshold, not a subscription). An evaluation and hobby budget, but a genuine programmable API at $0.

from openai import OpenAI
import os

client = OpenAI(
    api_key=os.environ["OPENROUTER_API_KEY"],
    base_url="https://openrouter.ai/api/v1",
)

resp = client.chat.completions.create(
    model="qwen/qwen3-coder",   # paid; the :free variant was retired
    messages=[
        {"role": "system", "content": "You are a concise coding assistant."},
        {"role": "user", "content": "Write a Python function that flattens an arbitrarily nested list."},
    ],
)
print(resp.choices[0].message.content)

Because it's OpenAI-compatible, the same code talks to any other OpenRouter model — swap the model string for DeepSeek, GLM, or Llama.

Way 2 — Self-host the open weights (Qwen3-Coder-Flash)

With Apache 2.0 weights, you can run Qwen3-Coder yourself with zero API keys and zero usage limits — free forever, the only cost is hardware. The realistic target is Flash (30B-A3B); the 480B flagship needs multi-GPU memory. The fastest path is Ollama:

# pull and run Qwen3-Coder-Flash locally (no key, no quota)
ollama pull qwen3-coder:30b
ollama run qwen3-coder:30b

Ollama exposes an OpenAI-compatible server at http://localhost:11434/v1, so the same client code points at your own machine:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")  # key ignored locally

resp = client.chat.completions.create(
    model="qwen3-coder:30b",
    messages=[{"role": "user", "content": "Refactor this loop into a list comprehension:n"
               "result = []nfor x in nums:n    if x % 2 == 0:n        result.append(x * x)"}],
)
print(resp.choices[0].message.content)

For higher throughput, run the same weights under vLLM or SGLang, or use llama.cpp with a quantized GGUF on Apple Silicon. Self-hosting is the only route with no rate limit.

Way 3 — Puter.js (free, no key, in the browser)

For a front-end demo with no server, Puter.js uses a "User Pays" model — the end user's Puter account covers the tiny inference cost, so you ship no API key and pay nothing:

<script src="https://js.puter.com/v2/"></script>
<script>
  puter.ai.chat(
    "Write a JavaScript debounce function with a leading-edge option.",
    { model: "qwen3-coder" }
  ).then(resp => {
    document.body.innerText = resp.message.content;
  });
</script>

Unbeatable for a static site or teaching example; for a backend service use OpenRouter or self-hosting. Bonus: ModelScope offers a free daily API allocation (quota changes — check current), and chat.qwen.ai is a free web chat for quick sanity checks.

What Happened to Qwen Code's Free Tier?

Qwen Code — Alibaba's open-source terminal coding agent — is still free and open source; what changed is authentication. Through 2025 a Qwen OAuth sign-in granted a free daily quota from Alibaba. Per the official auth docs, that Qwen OAuth free tier was discontinued on 2026-04-15. To run Qwen Code at no cost now, point it at a third-party provider via settings.json — a local Ollama/vLLM server, or a model still on OpenRouter's free list such as cohere/north-mini-code:free. The headline: the tool is still free; the Alibaba-hosted quota behind it is gone.

Pricing (When You Outgrow Free)

The official hosted API lives in Alibaba Cloud Model Studio (international brand for DashScope). Pricing is pay-as-you-go per million tokens and is tiered by input length — longer contexts cost more per token. Rough shape as of mid-2026; confirm live numbers in Model Studio's model list before budgeting:

Model (API name)ContextInput / 1MOutput / 1MBest for
qwen3-coder-plusup to 1Mfrom ~$1 (tiered)from ~$5 (tiered)Flagship agentic coding
qwen3-coder-flashup to 1Mmuch lower (tiered)much lower (tiered)Cheaper, fast coding tasks
qwen/qwen3-coder (OpenRouter)1Msee OpenRoutersee OpenRouterPaid tier, one key

Two things before you scale: the per-token rate steps up as input crosses thresholds (past 128K, again toward 1M), so price a huge context at the higher tier; and there's no standing free grant on the official API — for strictly $0, stay on OpenRouter's :free variant or self-host.

Your First API Call (OpenAI-Compatible)

Every hosted route is OpenAI-compatible — same request/response shapes as OpenAI Chat Completions. The two base URLs:

  • OpenRouter: https://openrouter.ai/api/v1, model qwen/qwen3-coder (paid — the :free variant was retired).
  • Alibaba Model Studio (international): https://dashscope-intl.aliyuncs.com/compatible-mode/v1, model qwen3-coder-plus. (Mainland: https://dashscope.aliyuncs.com/compatible-mode/v1.)
from openai import OpenAI
import os

client = OpenAI(
    api_key=os.environ["OPENROUTER_API_KEY"],
    base_url="https://openrouter.ai/api/v1",
)

response = client.chat.completions.create(
    model="qwen/qwen3-coder",
    messages=[
        {"role": "system", "content": "You are a senior Python engineer."},
        {"role": "user", "content": "Write a decorator that retries a function up to 3 times with exponential backoff."},
    ],
)
print(response.choices[0].message.content)

To hit Alibaba's endpoint instead, change two lines — swap base_url to the DashScope URL and use model="qwen3-coder-plus". Add stream=True to stream tokens for responsive coding UIs.

Qwen3-Coder in Cline, Aider & Qwen Code

Alibaba doesn't publish a first-party Anthropic-compatible endpoint, so the clean path is the OpenAI-compatible base URL. In Cline, choose "OpenAI Compatible," set base URL to https://openrouter.ai/api/v1, paste your key, enter model qwen/qwen3-coder — note there is no :free suffix any more, so this route bills per token. In Aider:

export OPENAI_API_BASE="https://openrouter.ai/api/v1"
export OPENAI_API_KEY="$OPENROUTER_API_KEY"
aider --model openai/qwen/qwen3-coder

Or use Alibaba's own Qwen Code CLI, purpose-built for this model — configure it against OpenRouter or a local Ollama server via settings.json:

npm install -g @qwen-code/qwen-code
# then run `qwen` and pick a provider via /auth
# for a $0 route, configure OpenRouter with cohere/north-mini-code:free

To drive Claude Code, route through an OpenAI-to-Anthropic proxy or a gateway like LiteLLM, since there's no first-party Anthropic endpoint.

Qwen3-Coder vs Other Free Coding Models

FeatureQwen3-CoderKimi K2DeepSeekGLM (Z.ai)
Free API pathOpenRouter :free + self-hostOpenRouter :free + CF Workers AILow-cost paid (promos)3 free Flash models
Card to start (free path)NoNoYesNo
Open weightsYes (Apache 2.0)Yes (Modified MIT)Yes (MIT)Partly
Flagship context256K–1M~256K128K~200K
SWE-bench Verified68.4%65.8%HighHigh (coding Flash)
Runs locally on one GPUYes (30B Flash)No (1T MoE)Distills onlyFlash variants
Anthropic-compatible endpointNo (use proxy)YesVia proxyYes

The honest read: for the strongest openly downloadable coding model reached via OpenRouter's free variant or your own hardware, Qwen3-Coder is the pick — the 30B Flash running locally is a genuinely $0, no-rate-limit coding agent. If you want a model that drops into Claude Code via a first-party Anthropic endpoint, Kimi K2 or GLM's Flash tier are smoother. For zero-asterisk free with a first-party key, GLM's three free Flash models win. To route between all of them behind one key, use OpenRouter — and for the lowest token price on a first-party API, DeepSeek.

Frequently Asked Questions

Is the Qwen3-Coder API free?

Yes, through several routes — just not the one that used to be easiest. Qwen's free OAuth quota for the Qwen Code CLI was discontinued on 2026-04-15. But Qwen3-Coder is Apache 2.0 open-weights, so you can still run it for $0 by self-hosting the 30B Flash model locally — Apache-2.0 weights with no revenue cap, checked on the model card. The licence imposes no limit; your hardware does. At 30.5B total parameters (3.3B activated) this is a server-class download, not a laptop one, or use Puter.js in the browser. What you can no longer do is call it free on OpenRouter — the :free variant is gone. Alibaba's own hosted API (Model Studio) is pay-as-you-go.

Did Qwen Code stop being free?

The Qwen Code CLI is still free and open source. What ended on 2026-04-15 was the free Qwen OAuth daily quota that used to power it. To run it at no cost now, configure it against OpenRouter's free Qwen3-Coder model or a local Ollama/vLLM server instead of Qwen OAuth.

Can I run Qwen3-Coder on my own computer?

Yes. The Qwen3-Coder-Flash (30B-A3B) variant runs on a single high-VRAM consumer GPU or Apple Silicon via Ollama (ollama run qwen3-coder:30b), vLLM, or llama.cpp. The 480B flagship needs multi-GPU memory. The weights are Apache 2.0, so local use is free and unlimited.

How good is Qwen3-Coder at coding?

Very good for an open model. The 480B flagship scores 68.4% on SWE-bench Verified (342 of 500 real GitHub tasks resolved), which Alibaba positions as comparable to Claude Sonnet 4. It's tuned specifically for agentic, tool-using workflows rather than single-shot answers.

Related Reads