Neon learns science in the laboratories that produce its data.
Periodic specialized a one-trillion-parameter model with its own diffraction experiments. Neon achieves 55.3% on an internal test, compared to 2.7% for its base model.
A diffractometer directs X-rays at a material reduced to powder. The result is a series of peaks whose positions and intensities provide information about the crystalline structures present. When a sample contains several overlapping phases, determining what was actually produced can take an experienced scientist several hours.
Periodic wants to assign part of this work to Neon, a one-trillion-parameter model specialized in analyzing experiments conducted in its own laboratories. The company says it outperforms GPT-6 Astra and Claude Fable 5.1 on FrontierXRD, an internal evaluation focused on X-ray diffraction measurements.
The announcement does not describe a foundation model trained from scratch. Periodic started with Kimi K2.6, an open-weight model that also has one trillion parameters. Its teams then subjected it to midtraining followed by reinforcement learning based on scientific and experimental data.
Neon remains an internal model. Periodic has released neither its weights nor its training corpus, and it does not offer an API through which the model could be tested independently. The base model is open, but the specialized version, its laboratory data, its scientific environment, and FrontierXRD are not available for download.
The project is part of a broader strategy. Since its launch in 2025, Periodic has been building high-throughput laboratories in Menlo Park that operate continuously. Their work focuses in particular on superconductors, magnetic materials, and semiconductor materials.
These facilities are not intended solely to discover new compounds. They must also produce the data needed to train the company’s models. Each series of experiments generates successful measurements, failures, anomalies, and synthesis conditions that are rarely available in scientific publications.
Periodic organizes materials discovery around three questions. The first is what would be useful to make based on predicted properties. The second is how to synthesize the proposed material. The third is what was actually produced and whether its properties match the intended objective.
Neon initially operates at this third stage. It analyzes diffraction results to identify the phases present in a sample and estimate whether the experiment produced the expected material. These conclusions can then change the hypothesis or synthesis procedure used in the next cycle.
This loop distinguishes the project from a scientific assistant that relies solely on papers, public databases, and simulations. The laboratory produces new observations, the model interprets them, and its findings can then help teams select the next experiments.
This process remains slower than learning in a fully digital environment. A physical experiment requires equipment, materials, energy, and staff. It can take several hours or several days, and its outcome is not always clear enough to provide the model with an automatic reward.
Periodic is therefore attempting to separate the pace of laboratory work from that of computation. While new syntheses are underway, computing resources can continue analyzing existing measurements, producing hypotheses, and improving procedures. Completed experiments then feed into subsequent cycles.
The first highlighted result concerns FrontierXRD. The evaluation contains 134 samples from Periodic’s laboratories that were selected for their difficulty. The accepted solutions contain an average of five phases, whose signatures may overlap in the measurements.
In such cases, a mathematical match is not enough. A phase may appear to fit the shape of the peaks while remaining inconsistent with the elements used, the temperature, the atmosphere, or the sample’s history. The analysis must combine the measurements with synthesis conditions, related experiments, scientific literature, simulations, and crystal structure databases.
Kimi K2.6 initially achieved a 2.7% success rate on FrontierXRD. After midtraining and reinforcement learning, Periodic reports a score of 55.3% for Neon. That represents an improvement of slightly more than twentyfold.
The company also says Neon outperforms GPT-6 Astra and Claude Fable 5.1 while reducing the cost of each analysis. This conclusion is limited to FrontierXRD and the environment built by Periodic. It does not mean Neon has stronger general capabilities in science, programming, or reasoning.
All the models in the comparison use the Periodic Harness. This environment gives them access to experimental context, internal databases, simulations, and the company’s scientific tools. It also organizes the long reasoning sequences and tool calls needed to examine a sample.
This choice provides the models with the same resources during the comparison, but it prevents the entire result from being attributed to Neon. Periodic itself acknowledges that its environment increases Opus 5’s XRD analysis success rate by a factor of 3.8 compared with a setup based on Claude Code, open databases, and commonly used diffraction software, at a similar cost.
The ranking therefore measures a system comprising the model, its context, its tools, and its orchestration. A company without the same internal databases or laboratory information should not expect to reproduce the reported performance with Neon alone.
Success is not determined by an automatically verifiable solution either. Periodic uses an ensemble of two judge models, Opus 5 and GPT-5.6 Sol, to assess whether the proposed phases are compatible with the measurements and make chemical sense.
To prepare this system, specialists with doctorates in materials-related fields annotated several thousand measurements. Each example was reviewed by three people. Periodic says the human specialists agreed with one another in 77.2% of cases, while the automated judges agreed with them 74.6% of the time. Agreement reached 84% when their decisions were compared with the specialists’ consensus.
These figures show that automated evaluation approaches the level of disagreement