New AI Models on WidelAI: Opus 5.5, GPT-6 & Gemini 3.8
Meet Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and Gemini 3.8 Flash on WidelAI, with verified specs, credit rates, routing guidance, and rollout considerations.
WidelAI Team
Making leading AI models accessible in one workspace
New AI models on WidelAI in September 2026 give teams four meaningfully different choices rather than four interchangeable names. Claude Opus 5.5 is tuned for sustained coding and knowledge work. GPT-6 Sol targets complex coding and agentic workflows. GPT-6 Luna brings the same large context envelope to focused, high-volume jobs at a much lower token price. Gemini 3.8 Flash combines Flash economics with long-horizon software engineering and autonomous-agent positioning.
All four are now represented in the WidelAI model catalog, but selecting well still requires more than sorting by provider or headline price. Direct provider prices and WidelAI credits are related measurements, not the same unit. Context limits describe capacity, not a recommendation to send everything. Reasoning controls also differ enough that a migration should preserve intent rather than mechanically copy parameters.
This release guide explains what changed, where each model fits, how the rates translate, and how to evaluate the bundle in the WidelAI chat workspace. For deeper analysis, use the companion comparisons of Claude Opus 5.5 and GPT-6 Sol, GPT-6 Sol and GPT-6 Luna, and the practical Gemini 3.8 Flash guide.
What Joined the WidelAI Model Catalog
The September bundle spans three providers and four workload profiles. Anthropic released Claude Opus 5.5, model ID claude-opus-5-5, on September 22, 2026. OpenAI released GPT-6 Sol and GPT-6 Luna, model IDs gpt-6-sol and gpt-6-luna, on the same date. Google released the generally available Gemini 3.8 Flash, model ID gemini-3.8-flash, on September 2, 2026.
| Model | Provider position | Context | Maximum output | WidelAI credits per 1K input / output |
|---|---|---|---|---|
| Claude Opus 5.5 | Long-running agentic coding and knowledge work | 1M | 128K | 0.80 / 4.00 |
| GPT-6 Sol | Complex coding and agentic workflows | 1,050,000 | 128K | 0.40 / 2.00 |
| GPT-6 Luna | Efficient focused and high-volume work | 1,050,000 | 128K | 0.02 / 0.10 |
| Gemini 3.8 Flash | Long-horizon engineering, agents, enterprise workflows | 1M | 64K | 0.15 / 0.75 currently |
These categories are starting points, not guaranteed outcomes. A model positioned for agentic coding can still be wasteful on a classification job, while an efficient model can outperform a more expensive option on a narrow prompt with a crisp acceptance test. Browse the live model catalog when choosing because availability and product presentation can evolve independently of an announcement.
GPT-6 Astra is not one of these additions. It predates this bundle and remains an existing frontier reference for harder end-to-end work. Keeping that distinction matters when evaluating release timing or planning a migration.
Four Models, Four Practical Roles
Claude Opus 5.5 for sustained depth
Claude Opus 5.5 is the choice to test when the work requires a durable mental model across a long repository investigation, a multi-stage implementation, or a knowledge task with many connected constraints. Anthropic describes adaptive thinking as always on. WidelAI exposes low, medium, high, extra, and max public effort choices; extra maps to the provider's xhigh setting. That gives users a visible way to trade time and tokens for more deliberation without pretending effort is a quality guarantee.
The model has a June 2026 reliable knowledge cutoff, a 1M-token context window, and a 128K maximum output. Those limits make large tasks possible, but disciplined retrieval still beats indiscriminate context dumping. Stable instructions, a repository map, and the files implicated by evidence usually produce a clearer run than an unfiltered project archive.
GPT-6 Sol for complex implementation
GPT-6 Sol is OpenAI's member of this group for complex coding and agentic workflows. Its six reasoning settings are none, low, medium, high, xhigh, and max, with medium as the default. Direct integrations should favor the Responses API when using built-in tools or function calling. Chat Completions supports function calling only when reasoning is none, a caveat that can break a migration if an application changes the model ID but retains a reasoning level.
GPT-6 Luna for scale
GPT-6 Luna carries the same 1,050,000-token context and 128K maximum output represented in the repository catalog, but its economic purpose is different. It is the most efficient GPT-6 option for focused and high-volume work. That makes it a strong candidate for extraction, tagging, templated transformations, short drafting, deterministic tool selection, and first-pass triage before escalation to Sol or Opus.
Luna supports the same six reasoning levels and the same Responses-versus-Chat-Completions caveat. Shared controls reduce integration friction, but they do not make output quality or token use identical.
Gemini 3.8 Flash for fast long-horizon work
Gemini 3.8 Flash is generally available and positioned for long-horizon software engineering, autonomous agents, and enterprise workflows. Thinking levels are low, medium, and high, with medium as the default. Google reports 54.9% on HLE-Verified; that is a vendor-reported result, not an independent head-to-head benchmark against the other three models.
Google also notes that the model may intentionally use more tokens on difficult tasks. Flash therefore describes a speed-and-cost family, not a promise that every hard answer will be short.
Provider Prices and WidelAI Credits
Direct provider prices are quoted per million tokens. WidelAI usage is quoted in credits per thousand tokens, where one credit equals $0.005. The conversion used for these rates is direct price per 1M multiplied by 0.2 to obtain credits per 1K. For example, a $4-per-million input price becomes 0.80 credits per thousand input tokens.
| Model | Direct input / output per 1M | Cache pricing per 1M | WidelAI input / output per 1K |
|---|---|---|---|
| Claude Opus 5.5 | $4 / $20 | $0.20 reads; $5 five-minute writes | 0.80 / 4.00 credits |
| GPT-6 Sol | $2 / $10 | $0.20 cached input; $2.50 writes | 0.40 / 2.00 credits |
| GPT-6 Luna | $0.10 / $0.50 | $0.01 cached input; $0.125 writes | 0.02 / 0.10 credits |
| Gemini 3.8 Flash | $0.75 / $3.75 introductory | Check the current provider pricing page | 0.15 / 0.75 credits currently |
Gemini's introductory direct rate runs through December 31, 2026. Standard pricing becomes $1.50 per million input tokens and $7.50 per million output tokens from January 1, 2027. The current WidelAI values reflect the introductory direct rates; check transparent pricing before forecasting a later workload.
Provider-specific service tiers complicate a direct invoice comparison. Anthropic reports an $8/$40 fast mode for Opus 5.5. OpenAI states that Batch and Flex cost 50% of Standard, fast mode costs 2× where applicable, and regional processing adds 10%. Those are direct-provider options and should not be presented as WidelAI credit modes unless the WidelAI product explicitly exposes them.
Context Windows Need a Budget, Not a Brag
All four models can accept very large prompts: 1M tokens for Opus 5.5 and Gemini 3.8 Flash, and 1,050,000 for Sol and Luna. Opus, Sol, and Luna allow up to 128K output; Gemini allows 64K. Capacity helps with large evidence sets, but a context window is not free storage and does not remove the need to prioritize information.
The GPT-6 models have an especially important threshold. Above 272K prompt tokens, the entire request is billed at 2× input and cache rates and 1.5× output rates. Crossing the line by a small amount can therefore change the economics of every token in that request. Before sending hundreds of thousands of tokens, retrieve only relevant files, remove repeated generated output, and summarize stable history with links to source artifacts.
A useful context plan has three layers:
- Put stable policy, definitions, and tool contracts in a reusable prefix.
- Retrieve evidence specific to the current task rather than attaching an entire corpus.
- Preserve authoritative outputs and decisions, while dropping speculative branches that no longer matter.
Long context is most valuable when the task genuinely depends on distant evidence. It is least valuable when it hides a vague prompt inside a large archive.
Reasoning and Tool Controls Are Not Portable Labels
Reasoning names differ across providers. Opus 5.5 uses adaptive thinking and WidelAI presents five public effort choices. Sol and Luna expose six choices, including none and max. Gemini 3.8 Flash has three thinking levels. “High” can mean a different budget and behavior on each family, so teams should compare completion quality, latency, output length, and total credits instead of assuming matching labels imply matching compute.
Tool interfaces also deserve explicit migration tests. OpenAI recommends Responses for built-in tools and function calling. Chat Completions function calling is limited to reasoning none for both GPT-6 models. Gemini migrations should replace thinking_budget with thinking_level, remove temperature, top_p, top_k, and candidate_count, and preserve interaction and function-call requirements. Opus integrations should account for always-on adaptive thinking rather than trying to disable the model's deliberation entirely.
A safe rollout treats a model change as a behavior change:
- Replay representative tool calls against a test environment.
- Validate argument schemas before execution.
- Require human confirmation for destructive or consequential actions.
- Keep credentials and filesystem access scoped to the minimum required.
- Record failures, retries, and tool-call loops, not only final-answer quality.
For security work, stay within systems you own or are authorized to assess. These models can assist with defensive review and remediation, but capability never supplies permission.
A Routing Strategy for Real Workloads
A multi-model workspace is most useful when routing rules are simple enough to follow. Begin with the least expensive model that reliably clears the acceptance bar, then escalate based on evidence rather than prestige.
| Workload | First model to test | Escalation signal |
|---|---|---|
| Repetitive extraction or classification | GPT-6 Luna | Missed schema fields or nuanced ambiguity |
| Fast implementation with tools | Gemini 3.8 Flash | Long-loop inconsistency or migration-specific limits |
| Complex multi-file coding agent | GPT-6 Sol | Root cause remains uncertain after evidence gathering |
| Long repository or knowledge investigation | Claude Opus 5.5 | Need a second provider perspective or different tool surface |
| Frontier cross-surface task | Existing GPT-6 Astra reference | Cost or latency is disproportionate to the task |
Use a fixed evaluation set before automating this router. Include easy, medium, and failure-prone cases. Score whether the task completed, whether tests or deterministic checks passed, how many human corrections were needed, elapsed time, WidelAI credits, and review defects. A model with a higher per-token rate can be cheaper if it succeeds once; a cheap model becomes expensive when loops multiply.
WidelAI's chat interface makes qualitative switching straightforward, but production routing still needs explicit logs and thresholds. Do not silently escalate indefinitely. Set a retry limit, a cost ceiling, and a rule for handing an uncertain result to a person.
Rollout Checklist for Teams
Establish a baseline
Capture current success rate, latency, output length, and cost on real tasks before changing models. Keep prompts, tools, source documents, and acceptance checks fixed. Without a baseline, a faster-looking answer can conceal more repair work.
Pilot by workload, not department
A single engineering team may need Luna for issue classification, Gemini for rapid implementation, Sol for complex agent runs, and Opus for deep diagnosis. Route by task shape rather than assigning one model to every person or application.
Verify integration-specific behavior
Test reasoning settings, function-call schemas, cancellation, retry behavior, context truncation, and cache eligibility. For GPT-6, test requests on both sides of the 272K threshold. For Gemini, test the January price change in budgets. For Opus, test adaptive-thinking behavior at each WidelAI effort option you plan to expose.
Keep review proportional to impact
Low-risk drafting may need sampling. Code changes need diffs and automated checks. Authentication, billing, infrastructure, data deletion, and security changes require explicit human review. Model sophistication does not reduce the consequence of an unchecked action.
Further Reading Across the Bundle
Each sibling guide goes deeper than an announcement can:
- Claude Opus 5.5 vs GPT-6 Sol examines coding, agents, cache economics, reasoning controls, and direct-price tradeoffs.
- GPT-6 Sol vs GPT-6 Luna explains when the 20× list-price gap is justified and when it is not.
- Gemini 3.8 Flash Guide covers API migration, thinking settings, pricing dates, evaluation, and rollout.
Primary references are Anthropic's Opus 5.5 announcement, pricing, and model overview; OpenAI's Sol documentation, Luna documentation, and family announcement; and Google's Gemini announcement, model guide, and pricing page. Source material is summarized and rephrased.
Conclusion: Treat the Bundle as a Routing Opportunity
The September 2026 additions matter because they broaden the useful middle of the model map. Claude Opus 5.5 offers sustained depth at $4/$20 direct rates. GPT-6 Sol offers a complex-agent tier at $2/$10. GPT-6 Luna makes focused GPT-6 work viable at $0.10/$0.50. Gemini 3.8 Flash combines a long-horizon brief with an introductory $0.75/$3.75 price through year-end.
The right response is not to crown one launch. Build a small representative evaluation, measure end-to-end success and total credits, and route each workload to the least expensive model that meets its quality bar. Recheck that decision when prompts grow beyond 272K tokens, when Gemini's introductory period ends, or when a task becomes consequential enough to justify a deeper model and stricter human review.
Do your best AI work in one place
Bring your next question, draft, file, or idea to WidelAI. Choose the model that fits the moment, keep your work together, and stay in control of privacy and usage.
Leading models, one workspace
Use powerful AI models without juggling separate tabs, accounts, or workflows.
Switch without starting over
Change models as your work evolves while keeping the conversation and context together.
The right model for every task
Choose speed for everyday work or deeper reasoning for complex questions and decisions.
Clear credits and model rates
See your balance, understand each model’s rate, and track usage from one place.
Bring your files and images
Work with documents and images alongside your prompts in the same focused experience.
Your work stays yours
Your data is encrypted in transit and at rest, and your content is never used to train AI models.