OpenAI's GPT-5.6 Luna Price Cut Drives Tenfold Usage Increase

Serge Bulaev

Serge Bulaev

OpenAI cut the price for its GPT-5.6 Luna model, which appears to have led to a tenfold increase in usage on OpenRouter. Reports suggest Luna quickly became the most used model by token volume, overtaking competitors like Anthropic's Opus 5 and Sonnet 5. The rise in usage may show that developers respond strongly to lower prices, using the model more for tasks like summarization and live inference. It is uncertain if Luna's popularity will last after the promotional pricing ends, but it currently remains one of the top models on OpenRouter. Analysts say these shifts in usage might keep happening if more price cuts occur.

OpenAI’s GPT-5.6 Luna Price Cut Drives Tenfold Usage Increase

A strategic price cut for OpenAI's GPT-5.6 Luna triggered a tenfold usage increase, dramatically boosting its market share on the OpenRouter platform. Speaking at a Goldman Sachs conference, OpenAI's finance chief Sarah Friar confirmed the price drop led to a ten-times jump in developer token consumption. The model quickly became the largest by volume on OpenRouter, a key hub for developers comparing AI models from various providers.

Industry analysts consider OpenRouter traffic a clear indicator of real-world model preference, as it isolates third-party developer activity. Data from mid-August confirms Friar's claim, showing Luna achieved a significant portion of token volume compared to its main competitors, Anthropic's Opus 5 and Sonnet 5. This event highlights how sensitive the AI model hierarchy is to pricing and suggests that developer demand for tokens is highly elastic.

Price Moves That Mattered

OpenAI's significant price reduction for its GPT-5.6 Luna model was the direct cause of the usage surge. By lowering the cost per token, the company made the model vastly more attractive to developers, who responded by increasing their consumption tenfold on platforms like OpenRouter for various tasks.

  • OpenAI implemented an 80% list-price reduction for Luna, setting the new rate at $0.20 per 1M input tokens and $1.20 per 1M output tokens, as detailed on its GPT-5.6 Luna - API Pricing & Benchmarks page.
  • OpenRouter had a promotion on eligible routes, but the exact pricing details for Luna during this promotional period are not confirmed by available sources.
  • The reduced price was consistently available across OpenAI, Azure EU, and Amazon Bedrock US, simplifying adoption for developer teams without requiring code modifications.

Usage Spike Quantified

Data published by OpenRouter confirms the dramatic growth, showing Luna's token share surged from 0.7% to 7.8% between July 27 and August 14. This increase represents a tenfold expansion in absolute token volume, as detailed in their analysis on GPT 5.6 Discounts & Jevons Paradox. For several consecutive days, Luna topped the daily rankings, outperforming models from Anthropic, DeepSeek, and Z AI.

A quick table distills the mid-August leaderboard derived from the public ranking API:

Rank Model Tokens per day (approx.)
1 GPT-5.6 Luna >2 trillion
2 Claude Opus 5 <1 trillion
3 Claude Sonnet 5 <1 trillion

Figures are rounded to the nearest trillion to reflect reporting thresholds; OpenRouter cautions that values represent routed share, not global usage.

Elastic Demand in Practice

Internal OpenAI data indicates the surge in token volume was substantial enough to more than compensate for the lower per-token rates, preventing a drop in revenue. This aligns with external analysis showing that as prices fall, developers increase experimentation, use larger context windows, and move workloads to the most cost-effective models. This behavior points to a price elasticity for developer APIs above 1.0, where total spending increases despite lower unit costs. Industry reports suggest this elasticity is significant for leading AI models.

Developers responding to an OpenRouter survey cited three practical ways they adapted to the lower Luna price:

  1. Increased Production Traffic: Developers replaced internal evaluation calls with live production traffic after confirming stable latency.
  2. Workload Optimization: Teams began splitting pipelines, using Luna for high-volume tasks like summarization while reserving premium models for complex reasoning.
  3. Shift to Live Inference: The negligible marginal cost per token prompted a shift away from caching toward real-time inference.

Competitive Implications

The long-term impact of this surge remains uncertain, especially as promotional pricing concludes. While early September data shows some usage shifting back to models from Tencent and DeepSeek, Luna maintains a top-three ranking in daily token volume. Analysts speculate that such rapid shifts in market share could become common if AI providers begin to compete by rotating monthly discounts.