GPT-6 Sol and Luna on Amazon Bedrock (September 2026): Routing Guide
Quick summary: On September 22, 2026 AWS made GPT-6 Sol and GPT-6 Luna generally available on Bedrock. OpenAI’s internal eval found Sol makes about half as many factual mistakes as GPT-5.6 Sol. On September 29 GPT-6.1 Sol followed, US Geo only. Here is the lane routing.
Key Takeaways
- On September 22, 2026 AWS made GPT-6 Sol and GPT-6 Luna generally available on Bedrock
- OpenAI’s internal eval found Sol makes about half as many factual mistakes as GPT-5
- 6 Sol
- On September 29 GPT-6
- 1 Sol followed, US Geo only
Table of Contents
AWS lifecycle notice (June 30, 2026) — Amazon Bedrock Agents Classic is in maintenance for new customers after July 30, 2026. Net-new agent builds should use Bedrock AgentCore. Full matrix: lifecycle roundup.
On September 22, 2026, AWS made GPT-6 Sol and GPT-6 Luna generally available on Amazon Bedrock. The Machine Learning Blog post positions them under GPT-6 Astra, which was already the top of the family. Sol is the daily model for recurring complex work. Luna is the high-volume model. On an internal OpenAI factuality evaluation, Sol makes roughly half as many mistakes as GPT-5.6 Sol.
This is a routing event on top of the July 30, 2026 GPT-5.6 price cuts. Do not flip production model IDs because the family number changed.
Successor: GPT-6.1 Sol (September 29, 2026)
What AWS changed: on September 29, 2026, GPT-6.1 Sol became available on Amazon Bedrock. Per the model card: text and image input, 1M context, 131,072 max output tokens, Standard tier only.
| Endpoint | Model ID | APIs |
|---|---|---|
| bedrock-mantle, us-east-1 only | openai.gpt-6.1-sol | Responses, Chat Completions (/openai/v1) |
| bedrock-runtime, US Geo CRIS only | us.openai.gpt-6.1-sol | Responses, Chat Completions, Converse, Invoke |
Per 1M tokens, 272K input tokens or fewer: $2.20 input, $2.75 cache write, $0.11 cache read, $11.00 output on both Mantle and US Geo. Above 272K input tokens the whole request moves to $4.40 input and $16.50 output. Input and output match GPT-6 Sol on US Geo. The cache read rate is half.
What the launch claims: OpenAI says GPT-6.1 Sol approaches GPT-6 Astra on demanding evaluations at roughly one-fifth of the cost per task. The ML Blog post says it matches Astra on DeepSWE v1.1 and beats the best GPT-6 Sol score by 6.4 percentage points. These are OpenAI’s numbers, not ours.
The catch:
- No global profile. GPT-6 Sol has
global.openai.gpt-6-sol; GPT-6.1 Sol does not, and there is no direct in-Region call. A lane that calls Sol fromeu-west-1orap-southeast-2through the global profile cannot move yet. - Prompt caching is unclear. The What’s New post mentions explicit prompt caching. The model card lists explicit prompt caching as not supported and publishes cache rates anyway. Do not budget on cache hits until a test request shows cached tokens in the usage block.
- Quota burns 10× on output. Each output token uses 10 tokens of quota. Check the account’s token-per-minute quota before moving a long-output coding lane.
Keep GPT-6 Sol as the prior Sol pin. Move a US coding lane to 6.1 only after a frozen-prompt bakeoff on cost per completed task. Keep Astra for the jobs where 6.1 still fails.
What each model is for
AWS and OpenAI describe three GPT-6 jobs, not one default:
| Model | Job | Bedrock status |
|---|---|---|
| GPT-6 Astra | Hardest end-to-end work: complex reasoning, coding, computer use, research, document creation | GA September 8, 2026. Model card: openai.gpt-6-astra. Mantle in us-west-2. Runtime profiles us.openai.gpt-6-astra and global.openai.gpt-6-astra. Context 1,050,000 tokens. Max output 128,000. |
| GPT-6 Sol | Recurring complex tasks and software development: implement, debug, review, refactor, analyze data, multistep tools | GA September 22, 2026. About half as many factual mistakes as GPT-5.6 Sol on OpenAI’s internal factuality eval. |
| GPT-6 Luna | High-volume extraction, summarization, classification, and routing, with adjustable reasoning effort | GA September 22, 2026. OpenAI reports better factual reliability than the prior Luna; the announcement does not publish a second numeric eval. |
Both Sol and Luna support up to 1M tokens of context on Bedrock (AWS). OpenAI’s API docs list a 1,050,000-token window for gpt-6-sol. Treat the Bedrock ceiling as the model card’s number once that card is in front of you — the two figures are close, and they are not the same sentence.
Opinionated take: new high-QPS classification, extraction, and routing go to GPT-6 Luna. New recurring coding and multistep tool lanes go to GPT-6 Sol. The hardest OpenAI jobs go to GPT-6 Astra. Keep GPT-5.6 Terra and Sol until a frozen-prompt bakeoff says the GPT-6 lane wins on quality and cost per completed task. Keep ZDR-default Claude work on Opus 5.5, not on an OpenAI model, when Guardrails composition and default zero data retention are the requirement.
Model IDs and endpoints
OpenAI’s Amazon Bedrock guide (checked September 23, 2026):
| Model | Mantle (in-region) | Runtime geo | Runtime global |
|---|---|---|---|
| Sol | openai.gpt-6-sol in us-east-1 | us.openai.gpt-6-sol | global.openai.gpt-6-sol |
| Luna | openai.gpt-6-luna in us-east-1 | us.openai.gpt-6-luna | global.openai.gpt-6-luna |
| Astra | openai.gpt-6-astra in us-west-2 | us.openai.gpt-6-astra | global.openai.gpt-6-astra |
Mantle base URL: https://bedrock-mantle.{region}.api.aws/openai/v1. Runtime base URL: https://bedrock-runtime.{region}.amazonaws.com/openai/v1, and the model name must be the inference profile — not the bare openai.gpt-6-sol ID. Hosted web search is Mantle-only. Background mode is Mantle-only and depends on data-retention settings. A region in the URL is not a data-residency guarantee; read the destination regions of the profile.
Context: OpenAI Python SDK, Mantle, us-east-1, Bedrock bearer token.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ['AWS_BEARER_TOKEN_BEDROCK'],
base_url='https://bedrock-mantle.us-east-1.api.aws/openai/v1',
)
response = client.responses.create(
model='openai.gpt-6-sol',
input='Summarize the failure modes in this migration plan and list what you could not confirm.',
)
print(response.output_text)For Runtime in the same region, point base_url at https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1 and set model to us.openai.gpt-6-sol or global.openai.gpt-6-sol. The guide’s older samples still show openai.gpt-5.6-sol in us-east-2. Changing only the model ID, and leaving the region, is the failure mode.
What broke (pattern, not a cited client) — A client that copied the GPT-5.6 Mantle sample, swapped in
openai.gpt-6-sol, and leftus-east-2in the base URL. Sol and Luna Mantle availability in the OpenAI guide is us-east-1. The call fails closed. Fix the region and the ID together, then confirm the live card.
Prompt caching and a three-model route
AWS says GPT-6 Sol and Luna support explicit prompt caching (GPT-6.1 Sol: see the caching conflict above). Mark stable instructions, tool definitions, policies, and extraction schemas so later calls process the new input. That matters when one request is classified by Luna, investigated by Sol, and escalated to Astra — each stage can reuse its own cached prefix instead of re-sending the policy pack.
Do not share one cache prefix across the three models. Cache identity is per model. A Luna checkpoint is not a Sol hit.
Prices: what is published, and what is not
AWS says Sol and Luna come at significantly lower API pricing than GPT-5.6 and does not print a rate card in the what’s-new post or the ML blog.
OpenAI’s API docs (first-party, not a Bedrock invoice) list, per 1M tokens:
| Model | Input | Cached input | Cache write | Output |
|---|---|---|---|---|
| GPT-6 Sol | $2.00 | $0.20 | $2.50 | $10.00 |
| GPT-6 Luna | $0.10 | $0.01 | $0.125 | $0.50 |
Prompts over 272K input tokens are 2× input and cache rates and 1.5× output for the whole request. Regional processing adds 10%. OpenAI’s Bedrock guide says commercial Bedrock pricing matches OpenAI direct pricing, and a region-specific Bedrock service is priced like OpenAI Regional processing. The Bedrock pricing page lists GPT-6 Sol and Luna as selectable models. The static extract on September 23, 2026 did not include their dollar rows. Confirm the selector before you re-forecast.
Update (September 30, 2026): the GPT-6 Sol model card now lists Bedrock Standard rates: $2.00 / $10.00 per 1M input/output on Global CRIS and $2.20 / $11.00 on Mantle and US Geo, with cache reads at $0.20 and $0.22. That matches the OpenAI list plus the 10% regional premium described above.
GPT-6 Astra does publish Bedrock Standard-tier rates on its model card. Short context (272K input tokens or fewer), per 1M tokens:
| Inference | Input | 30-minute cache write | Cache read | Output |
|---|---|---|---|---|
| Global CRIS | $10.00 | $12.50 | $1.00 | $50.00 |
| In-region and geo CRIS | $11.00 | $13.75 | $1.10 | $55.00 |
Long context (more than 272K input tokens) doubles input and cache rates and prices output at 1.5× ($75 global, $82.50 in-region). Priority and Flex are not supported for Astra. The card says in-region prices already include the 10% fee over OpenAI rates.
Previously published GPT-5.6 Bedrock on-demand pins ($0.22 / $1.32 Luna, $5.50 / $33 Sol) stay in our calculators. We are not replacing them with derived GPT-6 rates.
Data controls AWS actually named
- IAM for access, CloudTrail for every invocation, PrivateLink for VPC-only traffic.
- Inference runs on hardware-isolated infrastructure with zero-operator access.
- Inference data is not used for training, and you do not opt in to sharing it with OpenAI.
- Classifier-flagged abuse-detection traffic is retained by AWS for up to 30 days and processed programmatically. Zero data retention is a request through your account team.
That is a different default from Claude Opus 5.5, where AWS says ZDR is on by default on Bedrock. If the lane’s requirement is ZDR without an account-team exception, do not move it to GPT-6 because the coding eval looks better.
What to Do This Week
- Inventory production callers still pinned to
openai.gpt-5.6-luna,openai.gpt-5.6-terra, andopenai.gpt-5.6-sol. - Add Sol and Luna in a non-prod account. Smoke-test Mantle in us-east-1 and one Runtime profile (
us.orglobal.). - Pick one high-volume lane and one coding lane. Freeze the prompts. Score quality, latency, and cost per completed task against the GPT-5.6 pin.
- Turn on explicit prompt caching only for prefixes you can prove are identical across requests.
- For US callers, add GPT-6.1 Sol (
us.openai.gpt-6.1-sol) to the coding-lane bakeoff against GPT-6 Sol. Keep non-US callers onglobal.openai.gpt-6-sol. - Block a production default change until the Bedrock price selector shows a rate you are willing to budget, and until the bakeoff passes.
What This Post Doesn’t Cover
- A Bedrock dollar rate for Sol or Luna copied from the pricing page. The page lists the models; this pass did not capture dollar rows. OpenAI first-party rates are labeled as such above.
- A first-party FactualMinds bakeoff. The “half as many mistakes” figure is OpenAI’s internal factuality eval, not ours.
- GPT-6 Terra or a GPT-6.1 Luna. These announcements do not ship either.
- Codex or Managed Agents packaging. See the April OpenAI-on-Bedrock note.
Related reading
- GPT-5.6 Luna and Terra Bedrock Price Cuts (July 2026)
- Claude Opus 5.5 on AWS Bedrock (September 2026)
- OpenAI Models, Codex, and Managed Agents on Bedrock
- AWS Bedrock vs OpenAI API for Enterprise
- AWS Bedrock Cost Optimization: Token Budgets and Model Selection
Frequently asked questions
Are GPT-6 Sol and GPT-6 Luna generally available on Amazon Bedrock?
What are the Bedrock model IDs for GPT-6 Sol and Luna?
How much do GPT-6 Sol and Luna cost on Bedrock?
Should we replace GPT-5.6 Luna and Sol immediately?
What changed with GPT-6.1 Sol (September 29, 2026)?
Can I use GPT-6.1 Sol from outside the United States?
Does GPT-6 on Bedrock train on my prompts?

AWS Cloud Architect & AI Expert
AWS-certified cloud architect and AI expert with deep expertise in cloud migrations, cost optimization, and generative AI on AWS.



