Skip to content

Methodically built from peer-reviewed science.

Our engine assembles the applicable context, orchestrates a mixture of frontier models, and operates them in agents that think as scientists do.

It reasons across hours of compute and checks its own conclusions.

Closing the Empirical Loop: Autonomous AI Agents Conduct End-to-end Research With Human Participants

Wehr · Rideaux · Fox · Lightfoot · Tangen · Mattingley · Ehrhardt · Advanced Science (2026), e76675

§ 1

Made for doing science

01The engine

Multi-model orchestration

Explorer One works alongside frontier models, with every sub-task routed to whichever handles it best. Results are cross-checked across model families, removing single-model bias by construction.

OpenAI
Claude
Grok
Mistral
Gemini
02In-house infrastructure

The scientific tech stack

The stack operates the full scientific method autonomously. Verification checks each source and tests novelty against the field.

Dynamic retrieval

Live

The system retrieves and indexes literature at run time, never from a frozen training set.

DOI verification

Per-citation

The system resolves each reference to a live source before it enters an output.

Novelty assessment

Corpus-grounded

The system checks novelty against the published corpus, not the model’s prior.

Memory systems

Hours of context

The memory systems keep coherence across hundreds of agents and hours of compute.

03Scientific agents

Agents that think like scientists.

Multiple AI agents work together on research tasks. They spot patterns, break problems down, check their own work, and know when they're finished. The system adjusts in real time.

AgentField-IngestReads literature
AgentResearch-DesignCritiques design
AgentMethods & StatisticsAudits analysis
AgentPower-AnalysisSpawned · n=12
AgentInterpretive-RigourChecks claims
AgentCitation-HealthDOI verify · novelty
AgentAdversarialWrites case against
AgentSynthesisFinal output
AgentField-IngestReads literature
AgentResearch-DesignCritiques design
AgentMethods & StatisticsAudits analysis
AgentPower-AnalysisSpawned · n=12
AgentInterpretive-RigourChecks claims
AgentCitation-HealthDOI verify · novelty
AgentAdversarialWrites case against
AgentSynthesisFinal output
§ 2

The scientific method

The autonomous findings

Autonomous end-to-end scientific discovery

A single agentic system operates all steps in the scientific method.

  • Frame the question.
    Map the literature around it

  • Hypotheses, power, protocol. Pre-registered.

  • Run, log and verify at the source.

  • Statistics checked. Models stress-tested.

  • Draft, figures and a manuscript marked.

  • Reviewed, Calibre scored, novelty checked.

  • Journal fit, integrity screen, amplify.

Measuring the quality of science

Each manuscript review ends with a Calibre score, shown as a tier chip, but the number is only the entry point. Below it, the review gives the reasoning criterion by criterion: which parts are sound, which are not, and why.

The review ranks the issues by criticality, so the largest problems come first. Each issue has a specific fix that you can apply, which makes the score a map for the next revision.

Flawed Tier

Paper Score

Fundamental problems, stated openly.

0-9

Platinum Tier

Paper Score

Rigour at the level that demanding venues expect.

85-100

Gold Tier

Paper Score

Strong. Findings of the review are refinements.

70-84

Silver Tier

Paper Score

Sound core, named gaps to close before submission.

60-69

Bronze Tier

Paper Score

Mostly sound, with issues that weaken the claims.

50-59

Cobalt Tier

Paper Score

Notable weaknesses. Review is a map for the revision.

30-49

Tin Tier

Paper Score

Several central claims do not have sufficient support.

10-29

Flawed Tier

Paper Score

Fundamental problems, stated openly.

0-9

Platinum Tier

Paper Score

Rigour at the level that demanding venues expect.

85-100

Gold Tier

Paper Score

Strong. Findings of the review are refinements.

70-84

Unlocking scientific discovery.