Subquadratic presents SubQ, the first LLM with a fully subquadratic architecture
Subquadratic launches SubQ 1M-Preview, the first LLM with a subquadratic SSA architecture that cuts attention costs by 1,000x over 12 million tokens.
Miami-based startup Subquadratic is opening early access to SubQ 1M-Preview, touted as the first foundation language model built on an entirely subquadratic architecture. Where classic transformers see their computational cost grow quadratically with context length, SubQ relies on an attention mechanism dubbed SSA (Subquadratic Sparse Attention) that learns to retain only relevant relationships between tokens. Attention consumption is thus reduced by approximately 1,000 times over 12 million tokens compared to competing frontier models, with a speed 52 times faster than FlashAttention for 63% less computation.
On third-party verified benchmarks, SubQ 1M-Preview achieves 95% on RULER 128K (94.8% for Claude Opus 4.6), 81.8% on SWE-Bench Verified, and 65.9% on MRCR v2, which measures the ability to reason over long sequences. The research model, for its part, operates up to 12 million tokens, whereas competing models theoretically announce 1 million without always delivering on it.
Three products are opening in early access: a long-context API, SubQ Code (a command-line programming agent capable of loading an entire repository in a single pass), and SubQ Search (a deep search tool at conversational speed). The team comprises eleven researchers from Meta, Google, Oxford, Cambridge, ByteDance, Adobe, and Microsoft. The company has raised $29 million in seed funding, with investors including Justin Mateen, co-founder of Tinder, and several early backers of Anthropic, OpenAI, Stripe, and Brex.