This stabilization follows a period earlier in the year when the company’s CTO reported that Uber had already exhausted its entire 2024 AI budget.
The trend highlights a shift in how large enterprises manage AI infrastructure to prevent costs from ballooning alongside usage.
Uber achieved these savings by routing specific tasks to the most cost-effective models and implementing strict limits on "tokens"—the basic units of text or data that AI models process.
For instance, the company caps interactive sessions at 400,000 tokens and provides engineers with real-time cost tracking on their terminals to discourage wasteful spending.
The company is also moving away from high-cost strategies by experimenting with open-weight models, which are AI systems whose underlying data is publicly accessible and often cheaper to run.
CTO Praveen Neppalli noted that the company is shifting away from a period of unrestricted usage toward a more disciplined approach of matching specific tasks to the most efficient model available.
Beyond automated coding, the number of employees using these tools has more than quadrupled, with engineers now running over 30,000 agent-driven tasks daily.