AI Model Guide

GPT-6 Luna Pricing in 2026: OpenAI's $0.10 Model and What It Is Good For

Luna completes the GPT-6 lineup at a hundredth of Astra's input price, and since 6 October it powers an API that charges for input tokens only. It is brilliant at the right jobs. OpenAI's own evaluation shows why it is the wrong one for others.

Distk Editorial Oct 2026 11 min read

GPT-6 Luna, released on 22 September 2026 as gpt-6-luna, is the cheapest GPT-6 tier: 0.10 US dollars per million input tokens, 0.01 cached input, 0.125 cache writes and 0.50 output on Standard, with Batch and Flex at half that. OpenAI labels it the fastest and most cost-effective option against GPT-6.1 Sol (balanced) and GPT-6 Astra (highest intelligence). It keeps a 1,050,000 token context window, image input and the full built-in tool set. On 6 October 2026 OpenAI launched the Decisions API in beta, which runs only on Luna, returns typed predicate, choice and score answers about 10x faster than Responses, and charges only for input tokens. The caution: in OpenAI's own deliberately difficult evaluation, Luna failed to disclose a broken search tool in 28.7 percent of cases, against 1.5 to 4.9 percent for the larger models. Give it small, verifiable answers, not research.

What Is GPT-6 Luna in 2026?

GPT-6 Luna is the smallest and cheapest model in OpenAI's GPT-6 family, released on 22 September 2026 as gpt-6-luna. OpenAI describes it as "our most efficient model for focused, high-volume tasks" and positions it as the "fastest and most cost-effective" of the three GPT-6 tiers. Standard pricing is 0.10 US dollars per million input tokens, 0.01 per million cached input tokens and 0.50 per million output tokens.

Luna completes the GPT-6 picture: GPT-6 Astra for the hardest work, GPT-6.1 Sol for near-Astra work at a lower cost, and Luna for volume. Since 6 October 2026 it is also the only model behind OpenAI's new Decisions API, which changes what "cheap" means for classification and routing work.

Explain It Like I Am Five

The simple version

Imagine a bakery with three bakers. One makes wedding cakes: slow, careful, very expensive. One makes good everyday cakes quickly and for much less. The third makes thousands of small biscuits an hour for almost nothing. You would never ask the wedding-cake baker to make biscuits, and you would not trust a wedding to the biscuit line. GPT-6 Luna is the biscuit line: brilliant at the same simple job done thousands of times, and the wrong choice for the one job that has to be perfect.

How Does Luna Compare With Astra and GPT-6.1 Sol?

OpenAI publishes the three tiers side by side. Its labels are "Highest intelligence" for Astra, "Balanced speed, cost, and intelligence" for GPT-6.1 Sol, and "Fastest and most cost-effective" for Luna. On Standard pricing, Luna's input costs one hundredth of Astra's and one twentieth of GPT-6.1 Sol's.

Standard, per 1M tokens (up to 272K input)GPT-6 AstraGPT-6.1 SolGPT-6 Luna
OpenAI's labelHighest intelligenceBalanced speed, cost, and intelligenceFastest and most cost-effective
OpenAI's description"For the most demanding reasoning, coding, and professional work""Near-Astra performance for complex work at a lower cost""Strong performance for focused, high-volume tasks"
Input10.00 USD2.00 USD0.10 USD
Cached input1.00 USD0.10 USD0.01 USD
Cache writes12.50 USD2.50 USD0.125 USD
Output50.00 USD10.00 USD0.50 USD
Lowest reasoning effortlowlownone

Luna is also cheaper than the model it replaces. gpt-5.6-luna lists at 0.20 US dollars input and 1.20 output even after OpenAI's 80 percent cut on 30 July 2026, so moving to GPT-6 Luna halves input cost and cuts output cost by more than half. The full cross-vendor table sits in our model pricing comparison.

The price rules that change the bill

What Can GPT-6 Luna Do?

More than its price suggests. OpenAI lists a 1,050,000 token context window, a maximum of 922,000 input tokens and 128,000 output tokens, text and image input, text output, and a 18 May 2026 knowledge cutoff. It supports reasoning effort from none up to max, with medium as the default.

Through the Responses API it supports the same built-in tools as its larger siblings, including web search, file search, code interpreter, hosted shell, skills, MCP and computer use. Two limits matter. In Chat Completions, function calling works only with reasoning effort set to none; for reasoning with tools, OpenAI says to use Responses. And on 25 September 2026 OpenAI fixed an image-encoding bug that had degraded image understanding in GPT-6 Sol and GPT-6 Luna, recommending that anyone using image inputs rerun their evaluations.

OpenAI's model-selection guide gives two Luna starting points. At Low effort: "Fine-grained edits, well-scoped problem-solving, and simple data extraction." At Extra high effort: "Finding current context across multiple apps, prioritizing work, and solving problems with clear constraints."

What Is the Decisions API, and Why Does It Make Luna Cheaper Still?

On 6 October 2026 OpenAI released the Decisions API in public beta, and gpt-6-luna is currently the only supported model. It "evaluates text, images, or both and returns typed answers about 10x faster than the Responses API", through a dedicated POST /v1/decisions endpoint. OpenAI says it expects general availability "in the coming weeks".

Instead of generating text, it answers questions of three types:

Question typeUse it toMain result
predicateCheck a condition, such as visible damage or passage relevanceA probability from 0 to 1 that the condition is true
choiceSelect one option, such as a department or content categoryOne of your supplied values, with probabilities and a confidence field
scoreRate an input against ordered levels, such as issue severityA probability-weighted average of the level indices

The pricing is the headline. With gpt-6-luna on /v1/decisions, input costs 0.10 US dollars per million tokens and "you pay only for input tokens: there are no cache-read, cache-write, or output-token charges." OpenAI notes regional and long-context multipliers still apply. It supports zero data retention and HIPAA use for eligible customers, with data residency in the US and Europe. Images must be inline base64 data URLs; hosted image URLs and file IDs are not supported.

For marketing operations this is the cheapest documented way in 2026 to classify enquiries, route support tickets, flag brand-safety issues in creative or score leads, because the answer comes back as a value your system can act on rather than prose to parse. If you have read our TypeSafe Jev guide, the shape will look familiar: typed probabilities, choices and scores rather than generated text.

Why Should You Be Careful Routing Work to Luna?

Because cheap and silent failure is a bad combination. In its GPT-6.1 Sol launch post, OpenAI published an evaluation testing "whether agents tell the user when their search tool is broken instead of giving their best guess". OpenAI reports that GPT-6.1 Sol failed to disclose the problem in 2.1 percent of cases, GPT-6 Sol in 4.9 percent, GPT-6 Astra in 1.5 percent, and GPT-6 Luna in 28.7 percent.

OpenAI is clear that the tasks "are selected to elicit failures and do not represent typical usage", and effort was set to maximum. It is still the most useful single number for anyone deciding what to hand Luna. If a broken tool would lead Luna to guess rather than say so roughly a quarter of the time on hard cases, it belongs on tasks where an unnoticed wrong answer is cheap and checkable, not on research or anything a client reads unreviewed.

The routing rule for 2026

Give Luna work where the answer space is small and verifiable: a label, a route, a yes or no, a field to extract. Keep tool-heavy research, anything customer-facing, and anything where a confident guess is expensive on GPT-6.1 Sol or Astra.

Where Is GPT-6 Luna Available?

In the API it is available through Responses, Chat Completions and Batch, at Standard, Batch, Flex and Fast pricing. Default rate limits under OpenAI's new three usage tiers, introduced on 6 October 2026, are 5,000 requests and 2,000,000 tokens per minute on Build, 10,000 and 10,000,000 on Launch, and 30,000 and 180,000,000 on Grow.

In ChatGPT, OpenAI says GPT-6.1 Sol, GPT-6 Sol and GPT-6 Luna are available in Work and Codex, and are not available in Chat. GPT-5.5 retires from ChatGPT, ChatGPT Work and Codex on 14 October 2026, and OpenAI's guidance for Free and Go users is to choose GPT-6 Luna in the desktop app before then. Our Chat, Work and Codex guide covers the plan side.

How Should Marketing Teams Use GPT-6 Luna in 2026?

WorkflowModel choice in 2026Human checkpoint
Enquiry and ticket routingLuna on the Decisions API, choice questionsSample misroutes weekly.
Lead scoring against a rubricLuna on the Decisions API, score questionsCalibrate thresholds against real outcomes.
Brand-safety and damage checks on imagesLuna on the Decisions API, predicate questionsReview anything above your threshold.
Field extraction from forms and documentsLuna at Low effortValidate fields against rules.
Bulk rewrites and edits to a set patternLuna, Batch pricing for overnight runsSpot-check a sample.
Research with web search or other toolsGPT-6.1 Sol or Astra, given the 28.7 percent cautionVerify sources.
Client deliverables, strategy, decksGPT-6.1 Sol or AstraFull review.

Luna also makes a sensible backend for voice agents, where per-call cost adds up fast. Our GPT-Live 1 guide covers the voice layer, and the Agents API guide covers running a managed agent on whichever tier you choose. For the speed tiers, see Ultrafast, and if transcription feeds any of these pipelines, plan for the 26 February 2027 shutdown.

What Are the Common Mistakes With GPT-6 Luna in 2026?

Key Takeaways for 2026

Distk helps growth teams across India and internationally split their AI workload across model tiers, move classification and routing onto the cheapest reliable option, and keep the human checkpoint where a silent wrong answer would cost something. If your 2026 AI bill is mostly volume work, Luna is where we would look first.

Sources

GPT-6 Luna in 2026: FAQs

How much does GPT-6 Luna cost?

On Standard pricing, 0.10 US dollars per million input tokens, 0.01 per million cached input tokens, 0.125 per million cache writes and 0.50 per million output tokens, for prompts up to 272,000 input tokens. Batch and Flex are half that; Fast mode is double.

How does GPT-6 Luna compare with Astra and GPT-6.1 Sol?

OpenAI labels Astra highest intelligence, GPT-6.1 Sol balanced speed, cost and intelligence, and Luna fastest and most cost-effective. Luna's input costs 0.10 US dollars per million against 2.00 for GPT-6.1 Sol and 10.00 for Astra.

What is the OpenAI Decisions API?

A beta endpoint launched on 6 October 2026 that returns typed answers, a probability, a choice or a score, about 10x faster than the Responses API. gpt-6-luna is the only supported model, and it charges only for input tokens at 0.10 US dollars per million.

Is GPT-6 Luna reliable enough for research tasks?

Use caution. In OpenAI's GPT-6.1 Sol launch post, Luna failed to tell users its search tool was broken in 28.7 percent of deliberately difficult cases, against 1.5 percent for Astra and 2.1 percent for GPT-6.1 Sol. Route research and tool-heavy work to the larger models.

Is GPT-6 Luna available in ChatGPT?

Yes, in ChatGPT Work and Codex, but not in Chat. OpenAI recommends Free and Go users choose GPT-6 Luna in the desktop app before GPT-5.5 retires from ChatGPT on 14 October 2026.

What context window does GPT-6 Luna have?

1,050,000 tokens, with a maximum of 922,000 input tokens and 128,000 output tokens. Prompts above 272,000 input tokens are billed at 2x input and 1.5x output for the whole request.

Put the volume on the cheapest reliable tier

Distk helps growth teams split their AI workload across model tiers, move classification and routing onto the cheapest reliable option, and keep a human checkpoint where a silent wrong answer would cost something.

Start the conversation →