ESI / METHODS 001MOBSERVE
VERIFY
REVISE

METHODOLOGY · VERSION 1.0 · AUGUST 2026

How ESI reads the frontier.

The numbers on ESI are editorial indicators—not laboratory measurements. This page makes the judgment process visible, repeatable, and open to revision.

ESI tracks emerging capabilities by gathering public evidence, testing competing explanations, and assigning structured editorial judgments. Scores summarize that judgment so changes can be compared over time; they do not claim scientific precision or human equivalence.

Signal Pulse: evidence strength × trajectory.

Each 0–100 Signal Pulse combines five equally weighted questions: quality of evidence, repeatability, breadth across tasks or systems, real-world consequence, and rate of change. Editors score each question from 0 to 20 and add the results.

0–20Weak or anecdotal
21–40Early signal
41–60Material evidence
61–80Strong trajectory
81–100Frontier-defining

The ESI–1 through ESI–5 breakthrough label is a separate ordinal judgment: incremental, enabling, structural, transformative, or paradigm-level. It is not derived mechanically from Signal Pulse.

Current homepage values are the version-1 editorial baseline. Future revisions will include an “updated” date and a short change note when a score moves by five points or more.

Capability dimensions, scored separately.

The map avoids treating AGI as one finish line. Each dimension is evaluated against observable task performance, transfer to unfamiliar contexts, consistency under variation, and independence from human rescue.

  1. Collect representative evidence from public evaluations, system cards, demonstrations, and credible independent testing.
  2. Identify counterevidence, failure modes, and alternative explanations.
  3. Assign a 0–100 editorial estimate using the same anchored scale across updates.
  4. Record the date, evidence set, reviewer, and rationale for any material change.

“Now,” “Near,” and “Frontier” are horizon labels, not forecasts with fixed dates. Scores describe the strongest credible public systems ESI is tracking, not every deployed AI system.

Claims stay distinguishable from hypotheses.

ESI prioritizes primary sources, reproducible demonstrations, system documentation, and independent evaluations. Vendor claims and single demonstrations may identify a signal, but they receive less weight until replicated. Uncertainty is expressed through confidence labels and explicit caveats.

Update rhythm

The Radar and Development Map are reviewed monthly. ESI may update sooner after a major model release, replicated result, correction, or decisive counterexample.

Corrections

Material corrections should preserve the original publication date, identify what changed, and explain why. ESI does not silently rewrite a substantive claim.

A compact protocol for Action Lab work.

  1. State the task, success condition, model or tool version, and date.
  2. Run a single-agent baseline with the same information and time limit.
  3. Run the multi-agent or modified workflow at least three times.
  4. Preserve prompts, outputs, tool traces, failures, and human interventions.
  5. Compare accuracy, time, cost, recovery from error, and verification quality.
  6. Publish negative results and limitations alongside improvements.

ESI’s frameworks are working editorial instruments. They should improve as evidence, criticism, and better measurement practices become available.

OWNER: WARREN G · SYNTHETIC RESEARCH COLLABORATOR: CHATGPT · REVIEW STATUS: EDITORIAL FRAMEWORK, NOT PEER-REVIEWED