Vast-10M reaches ten million tokens without external memory
Voltropy is introducing Vast-10M, a family of LLMs with a native ten-million-token context window. Its Voltropy Scalable Attention technology modifies the attention mechanism of existing models and, according to the company's BEAM results, also improves reasoning at shorter context lengths.
Ten million tokens, actually tested
Advertising a ten-million-token context window does not show whether a model can still use information across that entire window. That is the question Voltropy is trying to address with Vast-10M.
The family includes three models. Vast-10M-Flash is based on DeepSeek V4.0 Flash, Vast-10M-Pro on DeepSeek V4.0 Pro, and Vast-10M-Medium on GLM-5.2. In each case, the original model's attention mechanism is replaced with Voltropy Scalable Attention, or VSA.
All three versions reach a native ten-million-token context without RAG or an external memory system. Voltropy then evaluates them from 100,000 to ten million tokens using BEAM, a benchmark designed to test reasoning and memory over extremely long contexts. Replacing attention in an existing model
VSA is designed to be retrofitted into existing Transformers rather than requiring an entirely new model architecture.
Voltropy's results show a less obvious effect: the modified models do not only improve at extreme context lengths. They consistently outperform their respective base models at 100K, 500K, and 1M tokens.
Vast-10M-Flash, for example, scores 52.66 on BEAM at 100K tokens, compared with 43.71 for DeepSeek V4.0 Flash. At 1M, the scores are 48.73 and 39.69. Vast-10M-Pro and Vast-10M-Medium also improve over their base models across all three shared context lengths.
Voltropy attributes those gains to VSA. The paper acknowledges additional training as a potential confounding variable when changing an existing model's attention mechanism, although the authors state that VSA required limited retraining for the model families used here. 48.73 against Fable 5.1 and GPT-6 Astra
At one million tokens, Vast-10M-Flash scores 48.73 on BEAM. Claude Fable 5.1 reaches 44.03 and GPT-6 Astra 49.75 in Voltropy's reported results.
That puts Vast-10M-Flash 4.70 points above Fable 5.1 and at 97.95% of Astra's score on this evaluation.
The comparison is also notable because DeepSeek V4.0 Flash, the base model, scores 39.69 at the same context length. The VSA version also outperforms DeepSeek V4.1 Flash at 100K, 500K, and 1M tokens.
These comparisons come from Voltropy's technical report and have not been independently evaluated. 40.20 points at ten million tokens
Performance declines as Vast-10M-Flash moves from one million to ten million tokens, but it retains 82.49% of its BEAM score: 48.73 at 1M versus 40.20 at 10M.
That 10M result is still slightly higher than the 39.69 achieved by its DeepSeek V4.0 Flash base model at only 1M tokens. It is also close to GLM-5.2's 40.49 score at 500K.
Vast-10M-Pro reaches 36.09 at 10M, while Vast-10M-Medium scores 32.74. Voltropy says Flash's 40.20 is the highest reported score for an unassisted language model on BEAM's 10M tier. The paper does not provide independent validation for that claim. A full benchmark run, with limitations
Voltropy says it ran the complete BEAM tests: 400 questions at 100K tokens, 700 at 500K, 700 at 1M, and 200 at 10M. GPT-4.1 mini was used as the official judge with temperature set to zero.
Direct comparisons between context tiers require some caution. The questions are not identical at every tier, although BEAM is designed to generate sets with comparable difficulty.
Fable 5.1 and GPT-6 Astra were only evaluated up to one million tokens in the report, based on the context windows considered by the authors. The paper therefore provides no Astra or Fable comparison against Vast-10M at the ten-million-token tier.
Early access to Vast-10M is now open through Voltropy.