The first flagship model under the unified SpaceXAI brand arrives focused on efficiency and automated code generation.
What Shipped
The model features a 1.5-trillion-parameter mixture-of-experts architecture on a new V9 base. It is trained on trillion of tokens of real Cursor agent-interaction data, signaling an integration that prioritizes practical developer application over sheer general capability. Review the corporate updates and background via SpaceX.
The Efficiency Pitch
The model scores 83.3% on Terminal-Bench 2.1 while using roughly a quarter of the output tokens Claude Opus 4.8 needs per solved SWE-Bench Pro task. It is priced at $2 per million input tokens and $6 per million output tokens, intentionally positioned as a high-volume cost-efficiency alternative in the coding agent market.
The Caveat That Matters
SpaceXAI self-disclosed that a Cursor codebase snapshot contaminated its training data, inflating its score on CursorBench specifically. This transparent self-disclosure highlights why engineering teams must maintain healthy skepticism regarding isolated benchmark superlatives across the industry. Validate real-world performance against internal repositories rather than relying solely on headline statistics. Explore integration details and IDE documentation at Cursor.
What It Means for You
This new flagship is worth testing for cost-sensitive, high-volume software engineering workloads. Account for the disclosed evaluation discrepancies and validate actual performance before finalizing your architecture decisions.