What Is GPT-6 Luna in 2026?
GPT-6 Luna is the smallest and cheapest model in OpenAI's GPT-6 family, released on 22 September 2026 as gpt-6-luna. OpenAI describes it as "our most efficient model for focused, high-volume tasks" and positions it as the "fastest and most cost-effective" of the three GPT-6 tiers. Standard pricing is 0.10 US dollars per million input tokens, 0.01 per million cached input tokens and 0.50 per million output tokens.
Luna completes the GPT-6 picture: GPT-6 Astra for the hardest work, GPT-6.1 Sol for near-Astra work at a lower cost, and Luna for volume. Since 6 October 2026 it is also the only model behind OpenAI's new Decisions API, which changes what "cheap" means for classification and routing work.
Explain It Like I Am Five
Imagine a bakery with three bakers. One makes wedding cakes: slow, careful, very expensive. One makes good everyday cakes quickly and for much less. The third makes thousands of small biscuits an hour for almost nothing. You would never ask the wedding-cake baker to make biscuits, and you would not trust a wedding to the biscuit line. GPT-6 Luna is the biscuit line: brilliant at the same simple job done thousands of times, and the wrong choice for the one job that has to be perfect.
How Does Luna Compare With Astra and GPT-6.1 Sol?
OpenAI publishes the three tiers side by side. Its labels are "Highest intelligence" for Astra, "Balanced speed, cost, and intelligence" for GPT-6.1 Sol, and "Fastest and most cost-effective" for Luna. On Standard pricing, Luna's input costs one hundredth of Astra's and one twentieth of GPT-6.1 Sol's.
| Standard, per 1M tokens (up to 272K input) | GPT-6 Astra | GPT-6.1 Sol | GPT-6 Luna |
|---|---|---|---|
| OpenAI's label | Highest intelligence | Balanced speed, cost, and intelligence | Fastest and most cost-effective |
| OpenAI's description | "For the most demanding reasoning, coding, and professional work" | "Near-Astra performance for complex work at a lower cost" | "Strong performance for focused, high-volume tasks" |
| Input | 10.00 USD | 2.00 USD | 0.10 USD |
| Cached input | 1.00 USD | 0.10 USD | 0.01 USD |
| Cache writes | 12.50 USD | 2.50 USD | 0.125 USD |
| Output | 50.00 USD | 10.00 USD | 0.50 USD |
| Lowest reasoning effort | low | low | none |
Luna is also cheaper than the model it replaces. gpt-5.6-luna lists at 0.20 US dollars input and 1.20 output even after OpenAI's 80 percent cut on 30 July 2026, so moving to GPT-6 Luna halves input cost and cuts output cost by more than half. The full cross-vendor table sits in our model pricing comparison.
The price rules that change the bill
- Long prompts: prompts with more than 272,000 input tokens are priced at 2x input and cache rates and 1.5x output for the whole request.
- Batch and Flex: 50 percent of Standard. Fast mode is 2x the applicable rates.
- Regional processing: adds a 10 percent premium where available. EU data residency is available with Standard, Fast, Flex and Batch.
- Caching: cached input is 10 percent of the uncached rate, and cache writes bill at 1.25x the uncached input rate.
What Can GPT-6 Luna Do?
More than its price suggests. OpenAI lists a 1,050,000 token context window, a maximum of 922,000 input tokens and 128,000 output tokens, text and image input, text output, and a 18 May 2026 knowledge cutoff. It supports reasoning effort from none up to max, with medium as the default.
Through the Responses API it supports the same built-in tools as its larger siblings, including web search, file search, code interpreter, hosted shell, skills, MCP and computer use. Two limits matter. In Chat Completions, function calling works only with reasoning effort set to none; for reasoning with tools, OpenAI says to use Responses. And on 25 September 2026 OpenAI fixed an image-encoding bug that had degraded image understanding in GPT-6 Sol and GPT-6 Luna, recommending that anyone using image inputs rerun their evaluations.
OpenAI's model-selection guide gives two Luna starting points. At Low effort: "Fine-grained edits, well-scoped problem-solving, and simple data extraction." At Extra high effort: "Finding current context across multiple apps, prioritizing work, and solving problems with clear constraints."
What Is the Decisions API, and Why Does It Make Luna Cheaper Still?
On 6 October 2026 OpenAI released the Decisions API in public beta, and gpt-6-luna is currently the only supported model. It "evaluates text, images, or both and returns typed answers about 10x faster than the Responses API", through a dedicated POST /v1/decisions endpoint. OpenAI says it expects general availability "in the coming weeks".
Instead of generating text, it answers questions of three types:
| Question type | Use it to | Main result |
|---|---|---|
predicate | Check a condition, such as visible damage or passage relevance | A probability from 0 to 1 that the condition is true |
choice | Select one option, such as a department or content category | One of your supplied values, with probabilities and a confidence field |
score | Rate an input against ordered levels, such as issue severity | A probability-weighted average of the level indices |
The pricing is the headline. With gpt-6-luna on /v1/decisions, input costs 0.10 US dollars per million tokens and "you pay only for input tokens: there are no cache-read, cache-write, or output-token charges." OpenAI notes regional and long-context multipliers still apply. It supports zero data retention and HIPAA use for eligible customers, with data residency in the US and Europe. Images must be inline base64 data URLs; hosted image URLs and file IDs are not supported.
For marketing operations this is the cheapest documented way in 2026 to classify enquiries, route support tickets, flag brand-safety issues in creative or score leads, because the answer comes back as a value your system can act on rather than prose to parse. If you have read our TypeSafe Jev guide, the shape will look familiar: typed probabilities, choices and scores rather than generated text.
Why Should You Be Careful Routing Work to Luna?
Because cheap and silent failure is a bad combination. In its GPT-6.1 Sol launch post, OpenAI published an evaluation testing "whether agents tell the user when their search tool is broken instead of giving their best guess". OpenAI reports that GPT-6.1 Sol failed to disclose the problem in 2.1 percent of cases, GPT-6 Sol in 4.9 percent, GPT-6 Astra in 1.5 percent, and GPT-6 Luna in 28.7 percent.
OpenAI is clear that the tasks "are selected to elicit failures and do not represent typical usage", and effort was set to maximum. It is still the most useful single number for anyone deciding what to hand Luna. If a broken tool would lead Luna to guess rather than say so roughly a quarter of the time on hard cases, it belongs on tasks where an unnoticed wrong answer is cheap and checkable, not on research or anything a client reads unreviewed.
Give Luna work where the answer space is small and verifiable: a label, a route, a yes or no, a field to extract. Keep tool-heavy research, anything customer-facing, and anything where a confident guess is expensive on GPT-6.1 Sol or Astra.
Where Is GPT-6 Luna Available?
In the API it is available through Responses, Chat Completions and Batch, at Standard, Batch, Flex and Fast pricing. Default rate limits under OpenAI's new three usage tiers, introduced on 6 October 2026, are 5,000 requests and 2,000,000 tokens per minute on Build, 10,000 and 10,000,000 on Launch, and 30,000 and 180,000,000 on Grow.
In ChatGPT, OpenAI says GPT-6.1 Sol, GPT-6 Sol and GPT-6 Luna are available in Work and Codex, and are not available in Chat. GPT-5.5 retires from ChatGPT, ChatGPT Work and Codex on 14 October 2026, and OpenAI's guidance for Free and Go users is to choose GPT-6 Luna in the desktop app before then. Our Chat, Work and Codex guide covers the plan side.
How Should Marketing Teams Use GPT-6 Luna in 2026?
| Workflow | Model choice in 2026 | Human checkpoint |
|---|---|---|
| Enquiry and ticket routing | Luna on the Decisions API, choice questions | Sample misroutes weekly. |
| Lead scoring against a rubric | Luna on the Decisions API, score questions | Calibrate thresholds against real outcomes. |
| Brand-safety and damage checks on images | Luna on the Decisions API, predicate questions | Review anything above your threshold. |
| Field extraction from forms and documents | Luna at Low effort | Validate fields against rules. |
| Bulk rewrites and edits to a set pattern | Luna, Batch pricing for overnight runs | Spot-check a sample. |
| Research with web search or other tools | GPT-6.1 Sol or Astra, given the 28.7 percent caution | Verify sources. |
| Client deliverables, strategy, decks | GPT-6.1 Sol or Astra | Full review. |
Luna also makes a sensible backend for voice agents, where per-call cost adds up fast. Our GPT-Live 1 guide covers the voice layer, and the Agents API guide covers running a managed agent on whichever tier you choose. For the speed tiers, see Ultrafast, and if transcription feeds any of these pipelines, plan for the 26 February 2027 shutdown.
What Are the Common Mistakes With GPT-6 Luna in 2026?
- Using it for tool-heavy research. OpenAI's own evaluation shows it hiding a broken search tool far more often than its siblings.
- Parsing generated text for a label. The Decisions API returns the label directly and charges input tokens only.
- Forgetting the long-prompt multiplier. Above 272,000 input tokens the whole request costs 2x input and 1.5x output.
- Calling functions in Chat Completions with reasoning on. That only works with reasoning effort set to
none. - Trusting image results from before 25 September. Rerun evaluations after OpenAI's image-encoding fix.
- Looking for it in ChatGPT Chat. It is in Work and Codex only.
- Treating the Decisions API as GA. It is a public beta.
Key Takeaways for 2026
- GPT-6 Luna launched on 22 September 2026 at 0.10 US dollars input, 0.01 cached input and 0.50 output per million tokens.
- It is OpenAI's "fastest and most cost-effective" GPT-6 tier, at a hundredth of Astra's input price and a twentieth of GPT-6.1 Sol's.
- It keeps a 1,050,000 token context window, image input and the full built-in tool set, including computer use.
- The Decisions API, in beta since 6 October 2026, runs only on Luna and charges only for input tokens at 0.10 US dollars per million.
- OpenAI reports Luna failed to disclose a broken search tool in 28.7 percent of deliberately difficult cases, against 1.5 to 4.9 percent for the larger models.
- Route small, verifiable answers to Luna; keep research and client-facing work on GPT-6.1 Sol or Astra.
Distk helps growth teams across India and internationally split their AI workload across model tiers, move classification and routing onto the cheapest reliable option, and keep the human checkpoint where a silent wrong answer would cost something. If your 2026 AI bill is mostly volume work, Luna is where we would look first.
Sources
- OpenAI, GPT-6 Luna model page.
- OpenAI, Using GPT-6.
- OpenAI, Model selection guide.
- OpenAI, Decisions API guide.
- OpenAI API pricing.
- OpenAI developer changelog, entries for 22 and 25 September and 6 October 2026.
- OpenAI, Codex models.
- OpenAI, Introducing GPT-6.1 Sol (broken search tool evaluation).