
AI shifts to "efficiencymaxxing" as inference costs loom large
AI teams are starting to focus more on "efficiencymaxxing," which means getting more output for each dollar spent, instead of just tracking how many tokens are used. This shift may be happening because running big AI models is getting more expensive as subsidies end, so companies need to use resources more wisely. Experts report that most of a model's energy use now comes from inference, and new methods appear to be making this step cheaper. Businesses are tracking new metrics like cost-per-million-tokens and ROI-per-token to watch spending. By using smarter routing and cheaper models, companies might keep quality high while reducing costs, and by 2026, about 40 percent of business apps may include specific AI agents to help with tasks.













