Every parameter count is a promise about memory and a distraction from what you pay. Ling 3.0 Flash VL carries 124 billion parameters and activates 5.5 billion of them per token, so most of the model sits still while a small slice does the arithmetic. Released on 10 September 2026 by InclusionAI and shipped open-weight under an MIT licence, it is one of the more consequential entries in this month’s run of latest ai models — and, depending on which public price list you read, either the cheapest multimodal model in the field or merely a cheap one. The two lists differ by more than a factor of three, and we will leave that disagreement standing rather than quietly pick a side. On OrcaRouter the latest ai models catalogue puts a rate beside every name, and a rate is only useful if you know which list it came from.
What 124B and 5.5B actually mean
A mixture-of-experts model publishes two numbers because it has two costs, and different people pay them. The total — 124 billion here — is what has to be resident, and it sizes the memory you need before the first token is generated. The active slice — 5.5 billion — is what runs per token, and it sizes the compute bill.
Vendors lead with the first number because it is larger; operators budget against the second because it recurs. Both appear as separate fields on the neutral third-party index’s page for this model, and the split is unusually wide even for this architecture family: roughly one part in twenty-two of the model participates in any single forward pass. The model is also a reasoning model, and the index records a reasoning-token budget of about 2,000 — tokens that are generated and therefore billed before the answer you wanted.
The two price lists do not agree
A third-party model catalogue that tracks this release lists it at $0.021 per million input tokens and $0.0616 per million output. The neutral index’s own page states $0.07 per million input and $0.22 per million output, annotating both as at the median of the models it tracks.
Those are not the same price: input differs by about 3.3×, output by about 3.6×. Neither figure is a typo we can find, and we are not going to average them. The gap most likely reflects that the two sources price different things. A catalogue rate is what one platform charges for its own serving of the weights, at whatever quantisation and hardware it chose; a leaderboard rate is what the index measured, or was quoted, for the endpoints it tests. When a model is open-weight and MIT-licensed anyone can serve it, so its price is not one number but a distribution.
The practical consequence is narrow: a claim that this model is the cheapest in its class is only true relative to a named source on a named date. Ask which.
What the neutral index measured
Where the index does have measurements it has a full set, and they describe a model with a pronounced shape. Its composite intelligence score is 24.57, recorded as measured rather than estimated; the rendered page rounds it to 25.
Underneath the composite sit two subscores that pull in opposite directions: analytical quality at an Elo of 904.49, interval 887.42 to 922.57, and presentation at 1072.4, interval 1051.99 to 1092.1. That is a gap of roughly 168 Elo points between how well it reasons and how well it writes, and the direction is the unusual one — a model whose output is better organised than its thinking suits drafting and structured extraction, and is a weaker bet where the answer depends on the reasoning being right. Its rubric pass rate is 0.249.
On speed the index reports a median of 144.84 tokens per second, against a field average of 132, with time to first chunk at 1.86 seconds.
One further measurement deserves its own sentence: over the index’s evaluation the model generated 160 million tokens against a median of 100 million. Output tokens are what you pay for, so verbosity converts directly into cost.
Where it degrades: context, tools and knowledge
The window is 262,144 tokens, and the index’s own breakdown shows what happens as you use it. Median output speed holds steady from 147.32 tokens per second on medium prompts to 144.84 on long, then drops to 128.79 on the 100K-token set. The latency figure is the more consequential one: median time to first chunk goes from 1.16 seconds on medium prompts to 8.67 on 100K, with end-to-end time rising to 28.08 seconds.
The capability picture is more uneven. Across the index’s individual evaluations the model scores 0.220 on a graduate-level question set, 0.442 on a scientific coding suite, 0.084 on a PDF-heavy task set, 0.02 on a physics-reasoning set, and 0 on the agentic terminal suite. On the knowledge measure it records −4.53, a negative value reflecting the index’s treatment of confident wrong answers.
Read those together with the presentation subscores and the shape is consistent: a fast, cheap, well-spoken model with a broad input surface and a narrow set of things it can reason through unaided. The defect is not in the release but in a deployment plan that assumes a low price implies a general substitute.
Licence, weights and what the input surface actually covers
The licence is MIT, and the index records the weights as available at InclusionAI’s Hugging Face repository. For a company that would otherwise rent this capability, the ceiling on cost becomes your own hardware rather than a vendor’s rate card, and the ceiling on availability your own operations rather than someone’s uptime page.
The modality mix is narrower than the marketing phrase suggests, and the index’s fields make it exact: text, images and video in; text only out. There is no audio input and no image or video generation.
Video input with text-only output is a defensible and quite specific product — describe what is in a clip, transcribe a screen recording, index footage into prose. It is not a media generation model, and “multimodal” without qualification will lead some buyers to expect one.
What to check before you route production traffic
Three checks, and none of them is a benchmark score.
First, settle the price question against your own serving path rather than either published list: pull the model, run your own token mix through it, and measure cost per accepted output. The 3.3× gap between the two public figures is a warning that per-token rates for open-weight models belong to the serving stack, not the weights.
Second, measure the verbosity tax directly: count output tokens per completed task on your traffic and multiply by the rate you actually pay.
Third, test the reasoning boundary rather than the average. The presentation-to-analytical gap and the zero on the agentic terminal suite point the same way — this model is strong where the task is to organise information and weak where the task is to derive it. Build two task sets and route accordingly.
Ling 3.0 Flash VL is not in OrcaRouter’s catalogue, so no rate of ours is quoted for it and nothing here implies we serve it. What makes it worth writing about is the shift it illustrates: when the weights are open and the licence permissive, the price of a model stops being a property of the model.
Reading a flash-tier release without overreading it
There is a temptation, with a release like this, to file it under a single adjective — cheap, or small, or open. The measurements above resist that.
The parameter split is the first thing to get right, because 124B and 5.5B describe different budgets and only one recurs. The price gap is the second, because a number quoted without a source is not a price. The capability profile is the third, because a composite score averages away the difference between the half of the work a model is good at and the half it is not.
What remains is a fairly clear picture: a 262,144-token window, three input modalities with text-only output, MIT-licensed weights, roughly 145 tokens per second, and a documented habit of saying more than the median model does. For extraction, description, transcription and drafting, those are good numbers at a low rate. For derivation and tool-driven autonomy, the index’s own subscores say to look elsewhere. Both statements come from the same page, and reading only one of them is how deployments go wrong.

