GPT-6 is a family of three models, and the choice between them plus a single reasoning.effort setting is now the main cost lever an engineering team has. OpenAI's guidance is blunt about the split. Use gpt-6-astra for the highest capability, gpt-6.1-sol for complex coding and professional work at a lower cost than Astra, and gpt-6-luna for efficient, repeatable, high-volume work (deployment checklist).
That matters because the price gap between the tiers is wide enough to change what you can afford to put in production. Astra runs $10.00 per million input tokens and $50.00 per million output tokens at short context, rising to $20.00 and $75.00 at long context. GPT-6.1 Sol is $2.00 in and $10.00 out at short context, $4.00 and $15.00 long (pricing). OpenAI's own announcement put GPT-6 Sol at half the price of GPT-5.6 Sol ($4 to $2 in, $20 to $10 out) and GPT-6 Luna at $0.10 in and $0.50 out, also half its predecessor (Introducing GPT-6 Sol and Luna).
A five-fold difference between Astra and Sol on input, and five-fold on output, means the model-selection decision is a budget decision made per workload, and it gets remade every time a new workload shows up.
What actually changed in the API
The Responses API is now the default path. OpenAI calls it the flagship API and the best place to access the newest model behavior, built-in tools, stateful workflows and agent features (deployment checklist). For GPT-6 this goes beyond style preference. Astra and GPT-6.1 Sol require Responses for tool calling, and GPT-6 Sol and GPT-6 Luna support function calling in Chat Completions only when reasoning_effort is set to "none".
The effort scale itself has grown. reasoning.effort guides how much the model thinks, and supported values are model-dependent, spanning none, minimal, low, medium, high, xhigh and max (reasoning models). Lower effort favours speed and fewer tokens. On GPT-6 Sol and GPT-6 Luna the supported set is none, low, medium (default), high, xhigh and max (GPT-6 Sol, GPT-6 Luna). Astra and GPT-6.1 Sol don't support none, and the GPT-6 guide notes that none and minimal are not supported there at all.
One more rule catches teams porting old code. When reasoning effort is not none, remove temperature and top_p from the request, because sampling parameters and reasoning effort don't coexist.
OpenAI's migration advice is to preserve your current model's workload role and effective reasoning effort where supported, then change one variable at a time. That is the right discipline. Swapping the model and dropping effort in the same deploy gives you a quality change you cannot attribute.
A walkthrough: three tiers in one pipeline
Take a document intake workload, a few hundred thousand PDFs a month that need classifying, extracting and, in a small number of cases, a judgement call. The temptation is to send everything to the best model. The cheaper and more defensible design routes by difficulty.
Classification and field extraction go to Luna at low effort. This is the high-volume, repeatable tier the docs describe, and at $0.10 per million input tokens it is the only tier where volume like that is comfortable.
# Illustrative: the cheap, high-volume tier
response = client.responses.create(
model="gpt-6-luna",
reasoning={"effort": "low"},
input="Classify this document and extract the invoice fields.",
)
Documents that fail a validation check, or come back with low-confidence fields, escalate to GPT-6.1 Sol at medium effort. OpenAI positions Sol for complex coding, computer use and professional work when you want near-Astra performance at lower cost, and recommends comparing it with Astra on your own tasks to judge the tradeoff (model selection). That comparison is the work. Nobody can tell you from a benchmark whether Sol is good enough for your contracts.
The narrow top tier, the exception queue a human would otherwise read, goes to Astra. Reserve xhigh and max effort for that queue, because effort buys reasoning tokens and reasoning tokens are output tokens at $50 to $75 per million.
For agent work, three tool features are worth knowing before you design the loop. Tool search loads deferred tool definitions at runtime so the model only imports what it needs, which avoids putting every tool definition in context up front and reduces token usage (tool search). Programmatic tool calling lets the model write and run JavaScript that orchestrates its tools, calling them in parallel, using loops and conditions, and keeping intermediate results in the hosted runtime (programmatic tool calling). Async tool calling lets GPT-6 keep reasoning, call other tools or answer independent parts of a request while your application runs a tool, with async: true on the tool.
Skills are the fourth. Agent Skills give an agent reusable instructions and supporting files for a task, usable with Responses API shell tools or in an Agents API sandbox (skills). If your team keeps pasting the same six paragraphs of house rules into prompts, that is a skill.
The cost controls you should turn on first
Prompt caching is the one that pays immediately. Reused prompt prefixes bill at the cached-input rate, discounted up to 95 percent, and cut time-to-first-token (prompt caching). For GPT-5.6 and later, cache writes cost 1.25 times the standard uncached input rate and subsequent reads cost 0.1 times it. The pricing table shows the same shape, with Astra cached input at $1.00 against $10.00 standard and Sol at $0.10 against $2.00.
The design consequence is to put stable content first. System instructions, tool definitions and policy text belong at the front of the prompt, with the variable user content at the end. Shuffle the order per request and you pay cache write prices forever.
Flex processing prices tokens at Batch API rates with additional prompt caching discounts, in exchange for slower responses and occasional resource unavailability (flex processing). OpenAI calls it ideal for non-production or lower priority tasks like model evaluations, data enrichment and asynchronous workloads. Your nightly backfill belongs there. Your customer-facing endpoint does not.
Running it on Databricks
If your data already lives in a lakehouse, the question is how to call GPT-6 without scattering API keys through notebooks. Databricks' answer has moved. Creating external model endpoints on Model Serving is now described as a legacy approach; for new workloads Databricks points to model provider services in Unity Gateway, which connect to external providers such as OpenAI and govern credentials, access, usage and cost with Unity Catalog (external models tutorial).
A model provider service is a Unity Catalog securable holding authentication and request configuration for an external provider (Model Provider Service API). You register the provider once, and many users and services route to it without ever handling the key (model providers external to Databricks). Creating one needs CREATE SERVICE on the schema plus USE CATALOG and USE SCHEMA, and the provider credentials themselves (create and manage providers).
The legacy path still works and still explains the shape of the thing. External models in Model Serving support openai as a provider, among anthropic, cohere, amazon-bedrock, google-cloud-vertex-ai and a custom option for OpenAI-compatible proxies (external models). Every external model served through Model Serving is queried with the OpenAI-compatible API, so one client works across providers. There is a constraint to watch, because Databricks returns an HTTP 4xx error if the provider doesn't support the model name you ask for, so confirm model availability before you wire a job to it.
Rate limits through AI Gateway support query-based (QPM) and token-based (TPM) limits, set per user, per group and endpoint-wide. For a GPT-6 rollout where one careless notebook at max effort can spend real money, the token-based limit is the control that matters. Set it before you hand out access, while the budget is still yours to protect.
If you are standing up governance around agents more broadly, the pattern we wrote about in Unity Catalog for AI agents in practice applies here too: the identity calling the model and the identity reading the data should be the same governed thing.
Three places GPT-6 will fight you
The first is anywhere you cannot change the API surface. If your integration is locked to Chat Completions with tools, Astra and GPT-6.1 Sol are out; you get function calling in Chat Completions only on Sol and Luna, and only at reasoning_effort: "none", which is the setting that turns reasoning off. Moving to Responses becomes a prerequisite for the whole thing.
The second is a tight latency budget combined with high effort. Higher effort means more reasoning tokens, which means more time and more output billing. A chat widget expected to start streaming in under a second is a low-effort workload or a Luna workload.
The third is any model you have tuned prompts against for a year. Sampling parameters are gone when effort is on, so prompts written to be nudged with temperature=0.2 need rewriting against the effort scale instead.
What I would tell a team starting this week
Pick one workload instead of the whole portfolio. Set the model to the tier you think is right and the effort to the default medium, then build the eval before you tune anything, because every later decision is a comparison and you need a baseline to compare against.
Then change one thing at a time: Astra to Sol, or medium to low, never both. Put the stable part of your prompt first so caching engages. Route the cheap tier to Luna and give the expensive tier a hard token limit. If the data is already in Databricks, register the provider in Unity Catalog instead of minting API keys per team, and set the TPM limit on day one.
More on how we approach this kind of work is on our AI page.
Frequently asked questions
Which GPT-6 model should I use?
Use gpt-6-astra for the highest capability, gpt-6.1-sol for complex coding and professional work at lower cost, and gpt-6-luna for cost-sensitive, high-volume workloads, per OpenAI's deployment checklist. OpenAI recommends comparing Sol against Astra on your own tasks to judge the quality and cost tradeoff.
What values does reasoning.effort accept?
Supported values are model-dependent and can include none, minimal, low, medium, high, xhigh and max, with medium the default on GPT-6 Sol and Luna (reasoning models). GPT-6 Astra and GPT-6.1 Sol don't support none or minimal.
How much does GPT-6 cost per million tokens?
At short context, gpt-6-astra is $10.00 input and $50.00 output, and gpt-6.1-sol is $2.00 input and $10.00 output; long context is $20.00/$75.00 and $4.00/$15.00 (pricing). Cached input drops to $1.00 and $0.10 respectively, and GPT-6 Luna is $0.10 input and $0.50 output (Introducing GPT-6 Sol and Luna).
Can I call GPT-6 from Databricks?
Yes, by registering OpenAI as a model provider service in Unity Gateway, which stores the credentials as a Unity Catalog securable and governs access, usage and cost (model providers external to Databricks). The older external model endpoints on Model Serving still work but Databricks now describes them as legacy for new workloads.