Scaling Test-Time Compute for Long-Horizon Scientific Reasoning
We show that allocating adaptive inference budgets across reasoning steps yields large gains on multi-stage scientific problem solving without retraining.
We show that allocating adaptive inference budgets across reasoning steps yields large gains on multi-stage scientific problem solving without retraining.
A latent diffusion transformer that generates novel protein backbones conditioned on functional motifs, with experimentally validated binders.
We introduce a memory-aware router that cuts activation footprint by 3x while matching dense baselines on standard language benchmarks.
Agents that write, execute and critique their own code improve steadily across iterations on a new benchmark of 1,200 research-engineering tasks.
SciFigure-Bench covers 40k figures from physics, chemistry and biology papers, revealing large gaps in current vision-language models.
Training directly on raw satellite observations, our model produces 10-day forecasts competitive with leading physics-based systems.
Neural noise models trained on calibration data reduce logical error rates by an order of magnitude on a 64-qubit superconducting device.
We study whether retrieval-grounded language models can flag statistical and design errors in submitted manuscripts, with surprising results.
An amortized reconstruction method that resolves heterogeneous molecular conformations from noisy cryo-EM images in minutes rather than days.