Skip to content

The Edge of the Cyber World See the latest

Apps

MiniMax M3.1 vs GPT-6 Luna vs DeepSeek V4.1 Flash: 3x Price Gap [2026]

Three new budget-tier language models shipped inside a five-day window in late September 2026, and none of them price themselves the same way. MiniMax dropped M3.1 Flash Preview into its coding agent on September 27 without publishing a single rate. OpenAI’s GPT-6 Luna, live since September 22, undercuts its own flagship Astra tier by roughly 100x on a per-token basis. DeepSeek V4.1 Flash, out since September 10, runs a peak-versus-off-peak pricing schedule that looks more like an airline fare chart than an API price list.

That gap matters for anyone routing high-volume traffic through a cheap model. A workload that costs $2.70 a day on DeepSeek V4.1 Flash during off-peak hours can cost double that at peak times, and a team building on MiniMax M3.1 Flash Preview still cannot forecast a monthly bill because the company has not released a rate card. This comparison breaks down what each model actually costs, how the three stack up on published benchmarks, and which one fits which job, using only pricing and benchmark figures that MiniMax, OpenAI, DeepSeek, or a verified third-party tracker has made public as of September 28, 2026.

None of these three models made the front page the way GPT-6 Astra or Claude Opus 5.5 did earlier in September. They matter for a quieter reason: this is the price tier where agentic AI actually runs in production. A support bot answering 50,000 tickets a month, a coding agent iterating through a pull request dozens of times before it passes review, a batch job summarizing a week of log files overnight, none of that work justifies flagship pricing. It runs on whichever model in this tier offers the best combination of speed, context length, and a predictable bill, and as of this week that answer is genuinely unsettled.

Why the Budget AI Model Tier Suddenly Matters

Frontier models got expensive fast this quarter. GPT-6 Astra, OpenAI’s top-end release from September 22, charges $10 per million input tokens and $50 per million output tokens, a price that makes sense for a handful of complex reasoning calls but breaks the budget the moment an agent starts making hundreds of tool calls per task. That is the gap the budget tier exists to fill, and three labs moved into it almost simultaneously.

DeepSeek shrank its key-value cache to roughly a quarter of the previous V4 Flash model’s footprint, which lowers the memory cost of holding long sessions open and lets the company pass some of that savings on through off-peak discounts. OpenAI split GPT-6 into three price points, with Luna sitting at the bottom specifically for high-throughput, low-complexity work. MiniMax took a different route entirely, shipping M3.1 Flash Preview straight into its MiniMax Code agent tool with a full 1-million-token context window but no published price, benchmark table, or model card at launch, according to Startup Fortune’s reporting on the release.

Every company chasing agentic workloads is now competing on tokens per dollar, not just raw intelligence scores. That shift explains why the three models compared here look so different from the GPT-6 Astra, Claude Opus 5.5, and Gemini 3.8 Flash flagships getting most of the September headlines. This is the tier developers actually pay for at scale, and it is the tier where the numbers stop being marketing copy and start being a real monthly invoice.

The September 2026 Release Timeline

The pace of releases this month is itself part of the story. DeepSeek moved first, shipping V4.1 Flash on September 10 and replacing both the earlier V4 Pro and the V4 Flash 0731 build in one step. OpenAI followed on September 22 with a double release, GPT-6 Sol and GPT-6 Luna landing on the same day as two different price points under the same GPT-6 umbrella that had already produced the flagship Astra tier. Anthropic released Claude Opus 5.5 that same day, September 22, putting three major labs on the same 24-hour news cycle. Google added Gemini 3.8 Flash TTS and a Flash-Lite TTS variant on September 23. MiniMax closed out the window on September 27 with M3.1 Flash Preview, and that same day a preview build of MiniMax Code went live carrying the new model’s five reasoning-effort levels, the fastest turnaround from model release to product integration of anything covered in this comparison.

That density of releases, five distinct model families in a span of 17 days, is unusual even by 2026 standards, and it is a big part of why pricing data for the newest entrants (MiniMax in particular) has not caught up with the launches themselves. Model cards, third-party benchmark runs, and aggregator listings all take time to populate after a release, and MiniMax’s is still in that gap a full day after launch as of this article’s publication.

MiniMax M3.1 Flash Preview: The Model With No Price Tag

MiniMax M3.1 Flash Preview launched on September 27, 2026, folded directly into MiniMax Code, the company’s agent-oriented coding tool. It carries a 1-million-token context window, matching the ceiling on the earlier MiniMax-M3 model, and MiniMax Code now exposes five distinct reasoning-effort levels for the model: low, medium, high, xhigh, and max, according to OrcaRouter’s breakdown of the launch.

What is missing is everything a buyer normally checks before committing production traffic. There is no per-token price published anywhere, no model card describing training data or safety evaluations, and no independent benchmark report covering coding, reasoning, or agentic tasks. There is also no way to switch the model’s thinking mode off, based on what OrcaRouter’s token-plan analysis found at launch. For a company that has previously shipped detailed spec sheets alongside its models, that silence stands out, and multiple outlets covering the release flagged it as unusual rather than routine.

None of that makes M3.1 Flash Preview unusable. It means treating it as an evaluation-only release rather than a production dependency until MiniMax publishes numbers. Teams already running DeepSeek’s V4.1 Flash or GPT-6 Luna in production have a documented cost basis to compare against. Anyone testing MiniMax right now is essentially flying without a fuel gauge.

The “Preview” tag in the model’s name is doing real work here. MiniMax has used preview labels before to gauge developer interest ahead of a wider rollout, and the absence of a rate card fits that pattern more than it suggests a permanently free or unpriced product. What is unusual is shipping a preview model directly inside a coding agent that developers might already be pointing at real repositories, rather than gating it behind a separate sandbox or waitlist. That choice puts more pressure on MiniMax to publish pricing quickly, since every day the model runs unpriced in a live tool is a day teams cannot responsibly forecast what switching to it, or staying on it, will eventually cost.

GPT-6 Luna: OpenAI’s Steep Discount Off Astra

GPT-6 Luna went live on September 22, 2026, alongside its sibling GPT-6 Sol, as the budget end of OpenAI’s newest model family. Standard pricing runs $0.10 per million input tokens and $0.50 per million output tokens. Cross the 272,000-token prompt threshold and those rates roughly double, to $0.20 input and $0.75 output, according to pricing tables published by MacRumors and corroborated by OpenRouter’s model listing.

Cached input drops to just $0.01 per million tokens, with cache writes at $0.125 per million, and batch-mode requests run even cheaper: $0.05 input and $0.25 output per million tokens, plus a $10 charge per 1,000 web-search calls when the model uses that tool. The context window sits at 1.05 million tokens, with a documented cap of 922,000 tokens on any single input and up to 128,000 tokens of output per response. OpenAI’s documentation lists six separate reasoning-effort settings for the model, giving developers a dial between speed and depth that MiniMax’s five-tier system loosely mirrors.

The number that puts Luna’s pricing in context is its distance from GPT-6 Astra. Astra charges $10 input and $50 output per million tokens. Luna charges $0.10 and $0.50. That is a 100x reduction on both sides of the ledger for a model built from the same family, and it is the clearest signal yet that OpenAI expects most agentic, high-call-volume traffic to run on the cheap tier rather than the flagship one. GPT-6 Sol and Luna’s combined launch already undercut Claude Opus 5.5 on price at the mid-tier, and Luna pushes that gap even further at the bottom.

DeepSeek V4.1 Flash: Peak-Hour vs Off-Peak Pricing

DeepSeek V4.1 Flash launched on September 10, 2026, replacing the earlier V4 Pro and V4 Flash 0731 models in the company’s lineup. Its pricing structure is the most unusual of the three: peak-hour rates run $0.30 per million input tokens and $1.20 per million output tokens, while off-peak rates drop to $0.15 and $0.60. Cache-hit pricing goes lower still, at $0.006 per million tokens at peak and $0.003 off-peak, based on figures reported by Requesty’s model pricing page.

Context runs to 1,048,576 tokens, close enough to the 1-million mark that DeepSeek describes it the same way MiniMax and OpenAI describe their own ceilings. Requesty lists a 384,000-token output cap for the model, while OpenRouter’s own listing shows completion tokens scaling up to the full 1,048,576-token context in some configurations. That discrepancy is worth flagging rather than smoothing over: different providers route the same underlying model through different serving configurations, and the effective output ceiling a developer sees can vary by access point.

Buyers should also expect a second layer of pricing confusion. DeepSeek’s own peak/off-peak schedule is not the only number circulating. OpenRouter lists V4.1 Flash access at $0.0295 per million input tokens and $0.60 per million output tokens, a rate noticeably below DeepSeek’s own off-peak input price. That gap reflects provider-specific routing and margin decisions rather than a conflicting official price, so anyone comparing quotes across aggregators should confirm which access path they are actually paying for before running the math on a production budget.

On the technical side, DeepSeek shrank the model’s key-value cache to roughly a quarter of the size used by the prior V4 Flash release, cutting the memory cost of holding long-context sessions open. The model supports both thinking and non-thinking modes, tool calling, JSON-structured output, and prompt caching, and raising the reasoning-effort setting from 25 to 100 consumes roughly 2.5 times more output tokens, according to DeepSeek’s own reporting cited by Dataconomy’s coverage of the launch.

Full Specs Comparison: 12 Data Points Side by Side

The table below lines up every publicly confirmed spec across the three models as of September 28, 2026. Where a company has not disclosed a figure, the table says so rather than guessing.

Spec MiniMax M3.1 Flash Preview GPT-6 Luna DeepSeek V4.1 Flash
Developer MiniMax OpenAI DeepSeek
Release date September 27, 2026 September 22, 2026 September 10, 2026
Context window 1,000,000 tokens 1,050,000 tokens 1,048,576 tokens
Max single input Not disclosed 922,000 tokens Not disclosed
Max output tokens Not disclosed 128,000 384,000 (reported by Requesty)
Standard input price /1M Not disclosed $0.10 $0.15 off-peak / $0.30 peak
Standard output price /1M Not disclosed $0.50 $0.60 off-peak / $1.20 peak
Cached input price /1M Not disclosed $0.01 $0.003 off-peak / $0.006 peak
Reasoning/effort levels 5 (low, medium, high, xhigh, max) 6 settings Adjustable effort (0-100 scale)
Thinking mode toggle Cannot be disabled Configurable Thinking and non-thinking modes
Tool calling / JSON output Coding-agent focused Supported Supported
Terminal-Bench 2.1 score Not published Not published 90.6
Primary integration MiniMax Code agent OpenAI API, batch endpoint DeepSeek API, OpenRouter, Requesty

Pricing Breakdown: What a Million Tokens Actually Costs

Raw per-token rates are hard to reason about until they are applied to a real workload. The table below models a mid-size agentic workload processing 10 million input tokens and 2 million output tokens in a single day, a realistic volume for a support bot or a coding agent running continuously across a small engineering team.

Model / tier Input cost (10M tokens) Output cost (2M tokens) Total daily cost
GPT-6 Luna, standard $1.00 $1.00 $2.00
GPT-6 Luna, above 272K context $2.00 $1.50 $3.50
DeepSeek V4.1 Flash, off-peak $1.50 $1.20 $2.70
DeepSeek V4.1 Flash, peak $3.00 $2.40 $5.40
DeepSeek V4.1 Flash, via OpenRouter $0.30 (approx.) $1.20 $1.50 (approx.)
MiniMax M3.1 Flash Preview Not calculable Not calculable No published rate

The headline gap is the input-token price between GPT-6 Luna’s standard rate and DeepSeek V4.1 Flash’s peak rate: $0.10 versus $0.30, a 3x difference on the exact same unit of work. Run that workload entirely during DeepSeek’s off-peak window and the gap narrows to 1.5x. Route it through OpenRouter’s DeepSeek listing instead of DeepSeek’s own API and the input price undercuts even Luna. The practical lesson is that DeepSeek V4.1 Flash is not one price, it is at least three, depending on time of day and which provider handles the request.

Scale the same math up to a month and the differences compound. At 10 million input and 2 million output tokens a day, GPT-6 Luna’s standard tier lands around $60 a month, DeepSeek V4.1 Flash off-peak lands around $81, and DeepSeek at peak hours climbs to roughly $162. Double the daily volume, a realistic jump for a growing product, and the gap between the cheapest and most expensive path through these three models widens from about $100 a month to well over $200. That is the kind of difference finance teams notice, and it is exactly why the peak/off-peak structure DeepSeek chose puts more scheduling burden on engineering teams than a flat rate does. A workload that can tolerate a few hours of delay is worth automating around that schedule. One that cannot, such as a live customer-facing chat feature, effectively pays the peak rate around the clock.

Benchmark Results From Three Independent Trackers

Benchmark data for this tier is thinner than for flagship models, but three separate trackers have published overlapping numbers. On the aggregate intelligence index maintained by BenchLM, GPT-6 Luna scores 66.59 and runs at a reported 162.9 tokens per second, the fastest of any model in this comparison. DeepSeek V4.1 Flash scores 39.5 on the same index, a few points ahead of Luna on raw intelligence but without a published throughput figure to match it against directly.

On agentic and coding-specific tests, DeepSeek V4.1 Flash is the only one of the three models with a substantial public benchmark record. It scores 90.6 on Terminal-Bench 2.1, ahead of the prior DeepSeek V4 Pro at 87.9 and edging out both Claude Opus 5 at 89.1 and GPT-5.6 Sol at 88.8 on that same test, according to benchmark tables published by DataCamp and corroborated by coverage from VentureBeat. The model also posts 74.2 on DeepSWE v1.1, 88.1 on CyberGym, and 54.8 on AutomationBench.

GPT-6 Luna has no independently verified coding-specific benchmark in circulation as of this writing. That is not unusual for a model positioned as the cheap, high-throughput tier rather than the reasoning flagship, but it does mean buyers evaluating Luna for agentic coding work are currently relying on OpenAI’s own effort-level documentation rather than third-party test results. MiniMax M3.1 Flash Preview has zero published benchmark numbers of any kind, which is the most significant gap in this entire comparison given that the model is already live in a production coding tool.

Context Windows and the Long-Context Pricing Cliff

All three models advertise context windows clustered around 1 million tokens: MiniMax at exactly 1,000,000, DeepSeek at 1,048,576, and OpenAI at 1,050,000. In practice that near-identical ceiling hides very different pricing behavior once a prompt gets long.

GPT-6 Luna applies a hard pricing cliff at 272,000 input tokens. Cross that line and both input and output rates roughly double, from $0.10/$0.50 to $0.20/$0.75 per million tokens. That threshold applies uniformly no matter when the request runs. DeepSeek V4.1 Flash has no equivalent token-count cliff, but it applies a time-based one instead: the same prompt costs twice as much if it happens to run during peak hours rather than off-peak. Neither structure is better or worse in the abstract, but they require different cost-monitoring approaches. A Luna-based system needs prompt-length alerts, while a DeepSeek-based system needs time-of-day-aware routing logic to avoid accidentally running batch jobs during expensive hours.

MiniMax M3.1 Flash Preview’s 1-million-token ceiling matches its predecessor MiniMax-M3, and since no pricing exists yet, there is no cliff to plan around at all, for better or worse.

Coding and Agentic Performance

MiniMax built M3.1 Flash Preview specifically for MiniMax Code, its agent-oriented coding tool, and gave it five reasoning-effort levels aimed squarely at multi-step coding tasks. Without a published benchmark, though, that positioning is a design intent rather than a proven result. Teams testing it for agentic coding should run their own evaluation suite before trusting it with anything beyond experimentation, a caution echoed by multiple outlets that covered the model’s quiet, spec-sheet-free launch.

A practical evaluation suite for this tier does not need to be elaborate. Pull 20 to 30 real tasks from your own backlog, the kind of pull requests, terminal commands, or ticket summaries the model would actually handle in production, and run each one through all three candidates at a matched reasoning-effort setting. Score correctness first, then compare token consumption per task, since a model that needs twice as many output tokens to reach the same answer erases most of a headline price advantage. That kind of side-by-side test is the only way to validate MiniMax’s claims today, given that no outside benchmark has done it yet.

DeepSeek V4.1 Flash currently has the strongest documented case for agentic coding work among the three. Its 90.6 score on Terminal-Bench 2.1, a benchmark built around real terminal and shell-based agent tasks, beats every other model named in that comparison table, including Claude Opus 5 and GPT-5.6 Sol. Its 74.2 on DeepSWE v1.1 backs that up with a second, independent measure of coding-agent competence.

GPT-6 Luna’s case for coding work rests on throughput and cost rather than benchmark scores. At 162.9 tokens per second and $0.10 input pricing, it is well suited to high-frequency, low-complexity coding tasks such as linting suggestions, commit-message generation, or lightweight code review comments, even without a headline coding benchmark to point to. Teams already running mixed workloads across providers often reach for a router layer like a LiteLLM-based fallback setup to send the easy calls to Luna and the harder agentic steps to DeepSeek V4.1 Flash automatically.

How the Wider Budget Tier Stacks Up

These three models do not exist in a vacuum. Grok 4.6 posts a reported 156.5 tokens per second on the same BenchLM tracker used for Luna, putting it in the same throughput class without matching Luna’s rock-bottom pricing. Gemini 3.8 Flash scores 40.9 on the intelligence index, a few points ahead of both DeepSeek V4.1 Flash and GPT-6 Luna, making it a reasonable fourth option for teams already inside Google’s ecosystem, though a separate lighter Flash-Lite variant does not yet have independently verified pricing or benchmark figures to compare directly.

Open-weight models are also crowding into this price band. Occamy-1.0, a 35-billion-parameter open-weight release from the Accio Team, scored 82.20 on the Claw-Eval co-work benchmark, edging out GPT-5.6 Sol at 81.80 and DeepSeek V4 Pro at 81.70 on that same test. Claw-Eval measures a different set of tasks than Terminal-Bench, so the scores are not directly interchangeable with the numbers earlier in this comparison, but the result signals that self-hosted, open-weight options are now competitive with mid-tier proprietary models on at least some benchmarks, which adds a fourth path for cost-conscious teams: running inference in-house entirely.

Grok 4.6 is worth a specific mention on price rather than just speed. It runs at the same $2 input and $6 output per million tokens as the newer Grok 4.7, according to pricing reported alongside xAI’s September 21 release notes. That places both Grok tiers well above every rate in this comparison’s main table, input pricing 6 to 20 times higher than GPT-6 Luna’s standard rate, depending on whether DeepSeek is running peak or off-peak. Grok remains a reasoning-first option rather than a budget one, and teams evaluating it alongside MiniMax, Luna, and DeepSeek V4.1 Flash should expect to pay a clear premium for it.

Availability, Access Paths, and What Isn’t Disclosed Yet

Access matters as much as price for teams deciding what to build on this week. GPT-6 Luna is available directly through OpenAI’s standard and batch API endpoints, with OpenRouter also listing it for teams that prefer a unified multi-model gateway. DeepSeek V4.1 Flash is available through DeepSeek’s own API with the peak/off-peak schedule described above, and independently through OpenRouter and Requesty, each running its own pricing and, in some cases, its own effective output ceiling. MiniMax M3.1 Flash Preview, by contrast, is reachable today mainly through MiniMax Code itself rather than a general-purpose, independently documented API that third-party developers can point arbitrary applications at.

None of the three companies has published detailed rate-limit figures, such as requests per minute or tokens per minute by account tier, in the sources reviewed for this comparison. That is normal for a launch week, since rate limits typically firm up as a provider observes real traffic patterns, but it is one more reason to treat any of these three as provisional for high-stakes production use until the surrounding documentation, not just the headline price, is fully in place. Teams that need firm rate-limit guarantees today should default to whichever provider’s status page and terms of service they can already point to, rather than assuming parity across all three.

5 Real-World Use Cases for This Model Tier

  • High-volume customer support triage. GPT-6 Luna’s $0.10/$0.50 standard pricing and 162.9 tokens-per-second throughput fit a support bot answering thousands of routine tickets a day, where individual call complexity is low and speed matters more than deep reasoning.
  • Agentic coding assistants working across full repositories. DeepSeek V4.1 Flash’s 90.6 Terminal-Bench score and 1,048,576-token context make it the strongest documented choice for an agent that needs to read, edit, and test code across many files in one session.
  • DevOps and terminal automation agents. The same Terminal-Bench result that favors DeepSeek for coding also applies directly to infrastructure agents executing shell commands, parsing logs, and chaining remediation steps.
  • Overnight batch summarization jobs. Scheduling large document or transcript summarization runs during DeepSeek V4.1 Flash’s off-peak window ($0.15/$0.60 per million tokens) costs roughly half what the same job would cost at peak hours, a straightforward scheduling win for non-urgent workloads.
  • Early-stage prototyping with a router safety net. Because MiniMax M3.1 Flash Preview has no published pricing or benchmarks, teams experimenting with it are better served pairing it with a fallback router that can reroute traffic to GPT-6 Luna or DeepSeek V4.1 Flash automatically if MiniMax pricing turns out to be uncompetitive once disclosed.

Which Model Should You Choose

If your workload is… Choose Why
High call volume, low complexity GPT-6 Luna Lowest standard input price and fastest reported throughput
Agentic coding across large codebases DeepSeek V4.1 Flash Highest published Terminal-Bench and DeepSWE scores
Non-urgent batch processing DeepSeek V4.1 Flash (off-peak) Roughly half the peak-hour cost for identical output
Extremely long single documents GPT-6 Luna 922,000-token max single input, the highest disclosed ceiling
Budget-critical production systems GPT-6 Luna or DeepSeek off-peak Both have transparent, published, predictable pricing
Experimental or research use only MiniMax M3.1 Flash Preview Capable spec sheet, but no price or benchmark to budget against yet

Migration Guide: Moving Workloads Between These Three APIs

Switching a production workload from one of these models to another is mostly a matter of remapping three things: the base URL and authentication, the reasoning-effort parameter, and the cost-monitoring logic. The steps below assume an existing OpenAI-compatible client setup, since DeepSeek and most aggregators expose OpenAI-compatible endpoints.

  • Benchmark your current per-request cost on the source model for at least 48 hours before migrating, so you have a real baseline rather than a theoretical one.
  • Confirm your longest real-world prompts fit inside the target model’s max single-input limit, especially when moving toward GPT-6 Luna’s 922,000-token cap.
  • Map reasoning-effort parameters across models: Luna’s six settings and DeepSeek’s 0-100 effort scale do not translate one-to-one, so test each target level against your actual task quality bar.
  • If moving to DeepSeek V4.1 Flash, build time-of-day-aware routing so batch and non-urgent jobs default to off-peak hours automatically.
  • Run your existing evaluation suite against MiniMax M3.1 Flash Preview manually if testing it, since no third-party benchmark exists yet to validate output quality against.
  • Set up a fallback router (LiteLLM is a common choice) so a pricing or availability change on any one provider does not take down the whole pipeline.
  • Re-run your cost baseline after migration and compare it against the pricing tables above to confirm the switch actually saved money at your real traffic pattern, not just on paper.

A minimal Python example switching between GPT-6 Luna and DeepSeek V4.1 Flash using the OpenAI SDK’s compatibility mode looks like this:

from openai import OpenAI

# GPT-6 Luna
luna_client = OpenAI(api_key="OPENAI_API_KEY")
luna_response = luna_client.chat.completions.create(
    model="gpt-6-luna",
    reasoning_effort="medium",
    messages=[{"role": "user", "content": "Summarize this ticket thread."}]
)

# DeepSeek V4.1 Flash (OpenAI-compatible endpoint)
deepseek_client = OpenAI(
    api_key="DEEPSEEK_API_KEY",
    base_url="https://api.deepseek.com/v1"
)
deepseek_response = deepseek_client.chat.completions.create(
    model="deepseek-v4.1-flash",
    messages=[{"role": "user", "content": "Summarize this ticket thread."}],
    extra_body={"reasoning_effort": 50}
)

MiniMax M3.1 Flash Preview is currently accessible through MiniMax Code rather than a standalone general-purpose API endpoint documented for third-party integration, so treat it as an in-tool evaluation rather than a drop-in migration target until MiniMax publishes broader access details.

Pros and Cons of Each Model

MiniMax M3.1 Flash Preview

  • Pro: full 1-million-token context window at launch
  • Pro: five reasoning-effort levels built for agentic coding
  • Con: no published pricing of any kind
  • Con: no model card or safety documentation
  • Con: no third-party benchmark results to verify quality claims

GPT-6 Luna

  • Pro: lowest standard input price of the three at $0.10 per million tokens
  • Pro: highest reported throughput at 162.9 tokens per second
  • Pro: transparent, fully published pricing including batch and cache rates
  • Con: pricing doubles above the 272,000-token prompt threshold
  • Con: no independently verified coding-specific benchmark available yet

DeepSeek V4.1 Flash

  • Pro: strongest published agentic and coding benchmark results in this comparison
  • Pro: off-peak pricing that undercuts GPT-6 Luna’s standard rate on output tokens
  • Pro: shrunk key-value cache lowers long-context memory costs
  • Con: peak-hour pricing roughly doubles the off-peak rate, complicating cost forecasts
  • Con: pricing varies meaningfully depending on whether you access it directly or through an aggregator

The Verdict: Which Budget Model Wins on Data

On the numbers published as of September 28, 2026, DeepSeek V4.1 Flash wins on documented capability. Its 90.6 Terminal-Bench 2.1 score beats every other model named in this comparison, including pricier options like Claude Opus 5 and GPT-5.6 Sol, and its off-peak pricing of $0.15/$0.60 per million tokens undercuts GPT-6 Luna’s output rate while staying close on input cost. For agentic coding and terminal-automation work specifically, it is the clearest choice among the three.

GPT-6 Luna wins on predictability and raw throughput. Its pricing is flat, fully disclosed, and the cheapest standard input rate of the three at $0.10 per million tokens, and its 162.9 tokens-per-second figure is the fastest reported speed in this comparison. For high-call-volume, lower-complexity work where budget forecasting matters more than benchmark supremacy, Luna is the safer bet.

MiniMax M3.1 Flash Preview remains the wildcard. A 1-million-token context window and a five-level reasoning-effort system suggest real capability aimed squarely at agentic coding, matching the same niche DeepSeek V4.1 Flash already serves with published proof. Until MiniMax releases pricing and benchmark data, though, it cannot be recommended for anything beyond side-by-side evaluation against the two models that already show their work.

The broader lesson from this comparison outlasts any single model’s specific rate. Budget-tier AI pricing in September 2026 is no longer a single number a company posts once and leaves alone. It is a live variable shaped by time of day, access path, prompt length, and how quickly a lab chooses to document what it shipped. Teams that build cost monitoring and multi-provider fallback into their AI infrastructure now will be far better positioned than teams that hardcode a single model and price into their system and assume it stays fixed. Given how fast this tier moved in a single month, betting on stability at these prices looks like the riskier choice.

Frequently Asked Questions

Is DeepSeek V4.1 Flash cheaper than GPT-6 Luna?
It depends on the hour. During DeepSeek’s off-peak window, its $0.15 input price is 50% higher than Luna’s $0.10 standard rate, but during peak hours DeepSeek’s $0.30 input price is 3x Luna’s rate. Output pricing follows the same pattern: DeepSeek off-peak ($0.60) undercuts nothing directly comparable, but DeepSeek peak ($1.20) is more than double Luna’s $0.50.

Why hasn’t MiniMax published pricing for M3.1 Flash Preview?
MiniMax has not given a public reason. Coverage of the September 27 launch from outlets tracking the release found no model card, benchmark table, or rate card attached, which is a departure from how MiniMax documented its earlier M3 model.

What is DeepSeek’s peak vs off-peak pricing exactly?
Peak-hour rates are $0.30 input and $1.20 output per million tokens. Off-peak rates are $0.15 input and $0.60 output per million tokens. Cache-hit input pricing drops further, to $0.006 at peak and $0.003 off-peak.

How does GPT-6 Luna compare to GPT-6 Astra on price?
Luna is roughly 100x cheaper than Astra on both input and output pricing: $0.10/$0.50 for Luna versus $10/$50 for Astra per million tokens.

Which model has the largest context window?
GPT-6 Luna’s 1,050,000-token window is marginally the largest of the three, ahead of DeepSeek V4.1 Flash at 1,048,576 tokens and MiniMax M3.1 Flash Preview at 1,000,000 tokens. The practical difference between all three is negligible.

Can I access these models through OpenRouter instead of the official APIs?
Yes for GPT-6 Luna and DeepSeek V4.1 Flash, both of which are listed on OpenRouter with their own pricing that can differ from the official API rates. MiniMax M3.1 Flash Preview is not yet broadly available through third-party aggregators as of this writing.

Is MiniMax M3.1 Flash Preview safe to use in production?
Treat it as evaluation-only for now. Without published pricing, a model card, or third-party benchmark results, there is no way to forecast cost or independently verify quality claims before committing production traffic to it.

Which model is best for agentic coding tasks specifically?
DeepSeek V4.1 Flash has the strongest documented case, with a 90.6 score on Terminal-Bench 2.1 and 74.2 on DeepSWE v1.1, both ahead of comparable mid-tier models. Grok 4.7 and Claude’s Opus 5.5 and Fable 5.1 tiers remain worth benchmarking against for teams whose budget tolerates a step up from this pure budget tier.

Related Coverage

Source: Tech Insider