What Is Gemini 3.5 Flash-Lite in 2026?
Gemini 3.5 Flash-Lite is Google's fastest and most cost-effective 3.5-class model, generally available in 2026 and positioned by Google DeepMind as best for low-latency, high-throughput agentic tasks. Google cites 350 output tokens per second according to the Artificial Analysis Index. It accepts text, image, video, audio and PDF input, returns text, supports a 1M-token input window and 64k output, and offers function calling, Search as a tool and computer use.
A version-number note before anything else, because it trips people up. The Flash-Lite tier is at 3.5 while the Flash tier is at 3.8 and the Pro tier is at 3.1 in preview. The number is the generation the model was trained in, not its rank in the lineup. Flash-Lite is always the smallest, fastest, cheapest tier, whatever number it carries, and Google states that on many agentic and coding evaluations 3.5 Flash-Lite even outperforms the older Gemini 3 Flash.
| Attribute | What Google published in 2026 | Why a marketing team should care |
|---|---|---|
| Status | General availability | Production-safe. |
| Positioning | Fastest, most cost-effective 3.5-class model; low latency, high throughput | This is the tier for the work you do hundreds of times a day. |
| Speed | 350 output tokens per second (Artificial Analysis Index, cited by Google) | Real-time classification and chat without visible lag. |
| Price | 0.30 USD per 1M input, 2.50 USD per 1M output, no caching | Up from 3.1 Flash-Lite on both sides; output up 67 percent. |
| Thinking levels | Selectable; more thinking for better reasoning and output quality | A per-workflow quality and cost dial. |
| Context | 1M input tokens, 64k output | Unusual at this price; bulk documents in one call. |
| Tool use | Function calling, Search as a tool, computer use | Computer use in the cheapest tier is the surprise. |
| Availability | Gemini App, Google AI Studio, Gemini Enterprise Agent Platform, Gemini API | Not listed for AI Mode or Antigravity. |
How Much Does Gemini 3.5 Flash-Lite Cost in 2026?
Gemini 3.5 Flash-Lite costs 0.30 US dollars per million input tokens and 2.50 US dollars per million output tokens in 2026, without caching, according to the pricing row in Google's own benchmark table. The same table lists Gemini 3.1 Flash-Lite at 0.25 and 1.50, GPT-5.4 mini at 0.75 and 4.50, and Claude Haiku 4.5 at 1.00 and 5.00. Google describes the result as a stronger price-to-performance ratio than 3.1 Flash-Lite, which is true, and it is also a price rise, which the page does not say in words.
| Model (per Google's table) | Input per 1M tokens | Output per 1M tokens | Versus 3.5 Flash-Lite |
|---|---|---|---|
| Gemini 3.5 Flash-Lite | 0.30 USD | 2.50 USD | Baseline |
| Gemini 3.1 Flash-Lite | 0.25 USD | 1.50 USD | 3.5 is 20 percent more on input, 67 percent more on output |
| GPT-5.4 mini | 0.75 USD | 4.50 USD | 2.5x on input, 1.8x on output |
| Claude Haiku 4.5 | 1.00 USD | 5.00 USD | 3.3x on input, 2x on output |
The output side is where the change lands. Output costs eight times input on 3.5 Flash-Lite, up from six times on 3.1 Flash-Lite. A classification workflow that reads a long input and returns a short label barely notices. A workflow that generates long product descriptions or translations at volume notices immediately. That asymmetry is the single most useful fact on the page for a budget owner in 2026.
A worked cost illustration for 2026
The table applies only the published rate cards to three hypothetical monthly volumes. The volumes are illustrative placeholders chosen to show the shape of the change, not measured figures from any deployment.
| Illustrative monthly volume | Gemini 3.1 Flash-Lite | Gemini 3.5 Flash-Lite | GPT-5.4 mini | Claude Haiku 4.5 |
|---|---|---|---|---|
| 50M input + 5M output (classification-heavy) | 20.00 USD | 27.50 USD | 60.00 USD | 75.00 USD |
| 50M input + 25M output (balanced) | 50.00 USD | 77.50 USD | 150.00 USD | 175.00 USD |
| 20M input + 40M output (generation-heavy) | 65.00 USD | 106.00 USD | 195.00 USD | 220.00 USD |
The costs above are arithmetic applied to the per-token rates in Google's published table at invented volumes. They exclude caching, thinking-level effects on output length, and any platform fees. They are not a forecast or an observed cost. Substitute your own measured token counts before using any of it in a budget.
What Do the Gemini 3.5 Flash-Lite Benchmarks Actually Say in 2026?
Google published a full comparison table, which is more than it did for 3.8 Flash. The comparators are Gemini 3.1 Flash-Lite, GPT-5.4 mini and Claude Haiku 4.5, so this is a small-model comparison and should be read as one. Every figure is Google's, run under Google's methodology, and the table is reproduced here as published.
| Benchmark (vendor-reported) | Gemini 3.5 Flash-Lite | Gemini 3.1 Flash-Lite | GPT-5.4 mini | Claude Haiku 4.5 | What it broadly measures |
|---|---|---|---|---|---|
| SWE-Bench Pro (Public) | 54.2% | 38.3% | 54.4% | 39.5% | Diverse agentic coding |
| Terminal-Bench 2.1 (Terminus-2 harness) | 54.0% | 31.0% | 59.2% | 44.2% | Agentic terminal coding |
| MLE-Bench | 39.2% | 22.0% | not listed | not listed | Machine learning engineering |
| GDPval-AA v2 (Elo) | 1140 | 642 | 1171 | 907 | Knowledge work |
| OSWorld-Verified | 74.0% | 54.3% | 72.1% | 50.7% | Agentic computer use |
| CharXiv Reasoning (no tools / with tools) | 74.5% / 76.5% | 73.2% / 75.6% | 80.3% / not listed | 61.7% / not listed | Reading complex charts |
| GDM-MRCR v2 8-needle, 128k average | 72.2% | 60.1% | 42.7% | 35.3% | Long-context retrieval |
| GDM-MRCR v2 8-needle, 1M pointwise | 21.3% | 12.3% | not listed | not listed | Very long-context retrieval |
Three readings for a business audience. First, the jump from 3.1 Flash-Lite is large everywhere, and on GDPval-AA knowledge work it is close to a doubling of the Elo score. That is the row that justifies the price rise for marketing work. Second, against GPT-5.4 mini it is a trade: Google leads on computer use and long context, OpenAI's small model leads narrowly on coding, knowledge work and chart reading. Third, the 1M-context retrieval score of 21.3 percent says what it says: the window is 1M tokens, but reliable recall at that length is not there yet at this tier. Use the big window for bulk throughput, not for needle-in-haystack accuracy.
What customers reported
Google quotes three partners. Ramp's head of applied AI says 3.5 Flash-Lite landed on the Pareto frontier of accuracy, latency and cost in its receipt extraction benchmark. Palo Alto Networks' principal AI engineer calls it a huge jump from 3.1 Flash-Lite and describes it as a responsive on-demand option for uninterrupted execution. Ashler's CTO cites agentic retrieval and tool use across fragmented infrastructure records while holding context across hundreds of deliverables. Receipt extraction is the one to remember: structured data out of messy multimodal input, at volume, where cost and latency decide.
How Should Marketing Teams Use Gemini 3.5 Flash-Lite in 2026?
Point it at everything that is high volume, well defined and verifiable. Google's own demonstrations show the pattern: receipt analysis and translation at scale, 25 web design options generated from one prompt, rapid iteration through multiple game builds, and a latency comparison against 3.5 Flash. The design demo is the most instructive because of how it was built: Gemini 3.6 Flash acted as the master agent and 3.5 Flash-Lite generated the 25 concepts underneath it. That orchestrator-and-worker split is the 2026 pattern for cost control.
The orchestrator-and-worker pattern
A larger model plans, decomposes and judges. A Lite model does the many parallel, repetitive sub-tasks. In a marketing pipeline that looks like a Flash or Pro model reading the brief and deciding what needs producing, then Flash-Lite generating twenty ad variants, classifying five hundred inbound enquiries, or translating a catalogue, with the larger model reviewing the result. The expensive model touches the work twice; the cheap model touches it hundreds of times.
| Marketing workflow | Fit for Gemini 3.5 Flash-Lite in 2026 | Thinking level and checkpoint |
|---|---|---|
| Inbound enquiry and lead classification | Strong. Bounded label set, measurable accuracy, latency matters. | Low thinking; sampled audit of misclassifications. |
| Receipt, invoice and form extraction | Strong. The Ramp use case, multimodal input to structured output. | Low to medium; validation rules on the output. |
| Catalogue and campaign translation | Strong on throughput; output tokens are the cost driver. | Medium; native-speaker review for customer-facing copy. |
| Ad copy variants at volume | Strong as a worker under an orchestrator. | Brand and claims review before launch. |
| Real-time chat and support triage | Strong. 350 tokens per second is the point. | Escalation path to a human on low confidence. |
| Chart and report reading | Workable; GPT-5.4 mini scores higher on Google's own table. | Spot-check numbers against source. |
| Long document Q&A at 1M tokens | Weak for precise recall; 21.3 percent on the 1M retrieval test. | Chunk it, or use a Flash-tier model. |
| Strategy, positioning, final quality review | Weak fit. Not what a Lite tier is for. | Route to Flash or Pro. |
What Do Thinking Levels Change for Cost and Quality in 2026?
Gemini 3.5 Flash-Lite lets the caller select how much thinking the model does before answering, and Google states that higher levels deliver improved reasoning and output quality. The practical effect for a budget is that thinking produces tokens, and tokens are billed. A classification task at the lowest level is the cheapest possible call; the same task at a high level costs more and is usually no more accurate, because the task did not need reasoning. The discipline in 2026 is to set the level per workflow and to measure accuracy at each level on your own evaluation set, rather than leaving a default in place across everything.
Where Does Gemini 3.5 Flash-Lite Fit in the Gemini Lineup in 2026?
At the bottom by design, and that is a compliment. The Flash tier's current model is Gemini 3.8 Flash, Google's workhorse for agents and document-heavy work. The Pro tier is Gemini 3.1 Pro in preview, with 3.5 Pro announced as coming soon. Flash-Lite exists so that the two tiers above it are never asked to do work that does not need them.
Two lineup details worth noting. Flash-Lite is listed for the Gemini App, AI Studio, the Enterprise Agent Platform and the API, but not for AI Mode or Antigravity, so it is a workflow model rather than a product surface. And it supports computer use, which means the orchestrator-and-worker pattern can extend to browser tasks: a Flash model deciding what to check, Flash-Lite clicking through fifty product pages to check it.
What Are the Common Mistakes to Avoid With Gemini 3.5 Flash-Lite in 2026?
- Reading "most cost-effective" as "cheaper than before." It is cheaper than the competition and more expensive than 3.1 Flash-Lite, especially on output. Re-run the budget.
- Ignoring the output asymmetry. Output costs eight times input. Generation-heavy workflows feel the price rise; classification workflows barely do.
- Trusting the 1M window for precise recall. 21.3 percent on the 1M retrieval test. Use the window for throughput, chunk for accuracy.
- Leaving thinking at a single default. Set it per workflow. Most Lite-tier tasks need the lowest level.
- Quoting Google's table as independent. It is vendor-run. GPT-5.4 mini and Haiku 4.5 are older small models, and OpenAI and Anthropic have both shipped since.
- Using a Lite model as the orchestrator. It is the worker. Let a larger model plan and judge.
- Skipping the human checkpoint because it is cheap. Low cost per call makes errors cheap to produce at scale, not cheap to fix.
Key Takeaways for 2026
Gemini 3.5 Flash-Lite is the model that will quietly run most of the volume in a Gemini-based marketing stack in 2026, and the release deserves the same budget scrutiny as a flagship.
- Generally available at 0.30 and 2.50 US dollars per million tokens, up from 3.1 Flash-Lite's 0.25 and 1.50, and well below GPT-5.4 mini and Claude Haiku 4.5.
- Google cites 350 output tokens per second, selectable thinking levels, 1M input, 64k output and computer use.
- Vendor-reported benchmarks show a large jump over 3.1 Flash-Lite, a lead over GPT-5.4 mini on computer use and long context, and a narrow trail on coding and knowledge work.
- Best fit: classification, extraction, translation, ad variants and real-time triage, as the worker under a larger orchestrator.
- Output costs eight times input. Generation-heavy workflows are where the price rise lands.
- The 1M window is for throughput; precise recall at that length scored 21.3 percent.
- Set thinking level per workflow, keep a fixed evaluation set, and keep the human checkpoint.
Distk helps growth teams across India and internationally design the orchestrator-and-worker split, instrument token usage per workflow, and put the review points where brand and revenue risk actually sit. If Gemini 3.5 Flash-Lite is carrying your volume in 2026, that design is where we start.