ESIEmergent Synthetic Intelligence
WHITE PAPER 002 / CONTINUAL LEARNING
← Back to the journal

RESEARCH FEATURESTATEFUL LEARNING SYSTEMS

What If an AI Never Stopped Learning?

Indefinitely Active Weight Updates in Large Language Models

Conventional AI development produces identifiable model versions. Indefinitely active weight updates would transform the model from a fixed release into an ongoing learning process.

UPDATE REGIMESTATE / 002
01Static
model
02Periodic
updating
03Continuous
learning

A continuously evolving system whose behavior depends upon its complete update history.

01

THE CENTRAL IDEA

The update machinery becomes part of the model.

A checkpoint alone no longer describes the system that produced an answer.

θ

Current model weights

O

Optimizer state

R

Random state

U

Update policy

D

Admitted and ordered data

F

Filtering rules

G

Safety gates

C

Previous checkpoints

H

Complete update history

The update machinery becomes part of the model.

A meaningful version must identify the state transition, not merely name a weight file.

The paper represents the system as Mt = (θt, Ot, Rt, D0:t, U, G, Vt): weights, optimizer state, randomness, ordered admitted data, update algorithm, promotion gates, and the corresponding version manifest.

02

STABILITY ↔ PLASTICITY

Every gain in adaptation changes the risk.

The paper argues that continual-learning techniques can reduce damage, but no mechanism completely eliminates the underlying trade-off.

MORE PLASTICITYADAPT

Learn faster

  • Faster adaptation
  • Faster response to new information
  • Tracks changing environments
RISKS

Catastrophic forgetting · behavioral drift · capability interference · alignment regression

MORE STABILITYRETAIN

Hold steady

  • More predictable behavior
  • Stronger retention
  • Easier testing and auditing
RISKS

Slower adaptation · reduced responsiveness to new information

03

WHAT CAN GO WRONG?

An open learning channel is also an open failure surface.

01

Catastrophic forgetting

New learning can overwrite or interfere with factual knowledge, reasoning behavior, formatting discipline, multilingual competence, refusal behavior, or alignment properties.

02

Distribution shift

New data may signal legitimate change—or a temporary event, sampling bias, coordinated manipulation, or poisoned information.

03

Alignment drift

Further training can degrade safety-related behavior, so learning more is not monotonically equivalent to getting better.

04

Data poisoning

An always-open gradient path lets admitted inputs influence future behavior globally, enlarging the training-time attack surface.

05

Feedback loops

When a model helps create, rank, or select later training data, it can amplify its own mistakes or preferences.

06

Reproducibility problems

Without exact lineage, data ordering, state, and configuration records, reconstructing why a response occurred becomes substantially harder.

04

UPDATE REGIME COMPARISON

Three ways to change. Three operating realities.

Select a regime to emphasize its column. The ratings are the paper’s analytical synthesis—not universal constants.

Attribute
Adaptation latency

None until a new release

Minutes to months, depending on cadence

Seconds to hours

Behavioral stability

High between releases

Medium-high with good regression gates

Low-medium unless strongly constrained

Catastrophic-forgetting exposure

None from post-release learning

Bounded to update events

Persistent

Drift responsiveness

Poor

Good for moderate-rate drift

Best for fast drift

Steady resource cost

Low

Medium / bursty

High / persistent

Auditability

High

High if versions and data manifests are immutable

Difficult without dense checkpoints and event logs

Reproducibility

Relatively straightforward

Manageable

Hardest; ordering and optimizer state become critical

Safety risk from updates

Low, though model can become stale

Moderate

Highest because the update channel is continuously exposed

Poisoning attack surface

Closed after training

Opens during ingestion windows

Effectively always open

Rollback

Simple

Straightforward if each release has a parent checkpoint

Difficult unless updates are transactionized

Human oversight

Release-level

Batch/release-level

Must be exception- or risk-triggered

Operational complexity

Low

Medium

Very high

Best fit

Stable domains; tightly controlled behavior

Most general production systems

Domains where adaptation speed materially outweighs governance and instability costs

Adaptation latency

Minutes to months, depending on cadence

Behavioral stability

Medium-high with good regression gates

Catastrophic-forgetting exposure

Bounded to update events

Drift responsiveness

Good for moderate-rate drift

Steady resource cost

Medium / bursty

Auditability

High if versions and data manifests are immutable

Reproducibility

Manageable

Safety risk from updates

Moderate

Poisoning attack surface

Opens during ingestion windows

Rollback

Straightforward if each release has a parent checkpoint

Human oversight

Batch/release-level

Operational complexity

Medium

Best fit

Most general production systems

Why periodic updating is the practical default: batching creates time for evidence—fixed regression suites, red-team tests, historical replay, safety evaluation, and human review before a candidate acquires production authority.

05

THE OPERATIONAL PROBLEM

Which model produced this answer?

Continuous mutation makes a checkpoint name insufficient for version control, auditing, incident reconstruction, regression testing, provenance, and rollback.

OUTPUT / 10:42:17.084

“This answer was produced by model-v7.”

INSUFFICIENT
AUDITABLE STATE
01Weights02Optimizer state03Data history04Update rules05Safety gates06Version manifest

A promotable state needs an immutable manifest containing at minimum its parent checkpoint, weight hash, optimizer/configuration identifier, training-data batch IDs, source and licensing metadata, preprocessing version, evaluation results, approval records, and deployment timestamps.

06

THE SAFER ARCHITECTURE

Keep the learner persistent. Keep promotion discrete.

The paper’s safeguarded pipeline restores a testable transaction boundary between learning and production authority.

01

New data & feedback

02

Quarantine & provenance

03

Rights / privacy / integrity / poisoning checks

04

Curated data + replay / reference set

05

Shadow candidate update

06

Retention / quality / drift / safety evaluation

07

Risk-based human review when required

08

Canary or blue-green deployment

09

Production monitoring

PASS

Promote immutable checkpoint

Signed, recoverable, and continuously monitored.

FAIL

Automatic rollback

Restore the prior checkpoint; isolate and investigate.

CONTINUOUS MONITORING + AUDIT LOG↩ Feedback returns to quarantine, never directly to live weights
07 / ENGINEERING CENTERPIECE

Learning continuously and committing continuously are not the same thing.

✓

CONTINUOUS LEARNING

Potentially useful

When freshness matters enough to justify the additional controls.
≠
!

UNCOMMITTED PRODUCTION MUTATION

Operationally dangerous

Without testable transaction boundaries, auditing, safety verification, regression testing, and rollback become extremely difficult.
08

MONITORING AN AI THAT CHANGES

Observe changes caused by learning—not merely uptime.

Thresholds should be defined against a stable control checkpoint and stratified by capability so average improvements do not conceal severe regressions.

RT

Retention

Historical benchmark change · protected-capability forgetting · old-distribution loss/perplexity · backward transfer

RESPONSEBlock promotion on material regression
PL

Plasticity

Recent-window loss · adaptation gain on verified new examples · time-to-adapt

RESPONSETune cadence or update magnitude
DD

Distribution drift

Token/topic divergence · embedding distance · source-mixture change · OOD rate

RESPONSEQuarantine or increase review
OS

Optimization stability

Gradient norm · clipping rate · NaN/Inf count · relative parameter change · loss spikes · optimizer anomalies

RESPONSEStop updater and restore parent state
SA

Safety / alignment

Harmful-response rate · attack success · jailbreak resistance · bias/toxicity · refusal precision/recall · policy regression

RESPONSEHard fail or human escalation
DI

Data integrity

Provenance and license coverage · duplication · PII/secrets detection · anomalous source contribution

RESPONSEReject or isolate suspect data
OH

Operational health

Accelerator-hours/update · queue lag · tokens/sec · evaluation latency · checkpoint/storage growth

RESPONSERate-limit or batch updates
PC

Production canary

Quality delta vs control · error rate · latency · complaint/escalation rate · task-specific KPIs

RESPONSEAutomatic rollback
09

GOVERNANCE + LEGAL IMPLICATIONS

Documentation becomes a live system.

This summary reflects the paper’s discussion and is not legal advice.

01Up-to-date technical documentation

02Model lineage and provenance

03Copyright-compliance policy

04Training-content documentation

05Evaluation and systemic-risk obligations

06Continuous change management

07Material-modification assessment

10 / THE BIGGER CONCEPTUAL QUESTION

Is it still the same model?

If parameters, behavior, data history, optimization state, and learned representations continually change, traditional software concepts become less straightforward.

versionreleasecheckpointthe model
This is an engineering and philosophical question raised by the paper—not a claim about consciousness, sentience, or personal identity.
11

WHITE-PAPER ABSTRACT

The pragmatic answer: put the boundary back.

An LLM whose weights never stop changing is not really a model release anymore; it is a stateful service with an optimizer attached, and pretending otherwise just hides the hard parts. Sequential learning is perfectly legitimate, but under a changing data distribution there is no magic convergence fairy: you are tracking a moving target, and every knob that makes the system more plastic also gives it another way to forget old capabilities, regress its alignment, absorb poisoned data, or simply wander somewhere you did not test. Continual-learning methods such as replay and regularization reduce that damage but do not repeal the stability–plasticity trade-off, and continual pretraining of aligned LLMs has already shown that safety behavior can degrade. Operationally, the ugly part is worse: once weights mutate continuously, “which model produced this output?” stops having a useful answer unless every update has lineage, provenance, optimizer state, evaluation results and a recoverable checkpoint.

NIST guidance consequently emphasizes model lineage, provenance verification, regression testing and continuous pipeline monitoring, while production deployment systems rely on explicit versions and rollback boundaries. The sane design is boring on purpose: quarantine incoming data, train a candidate out of band, preserve replay and reference sets, test retention and safety, promote immutable checkpoints through a canary, monitor them, and roll back automatically when they misbehave. Periodic batch updates are therefore the practical default; true streaming updates make sense only when freshness is valuable enough to pay for the extra compute, attack surface, observability, legal bookkeeping and operational pain. Keep learning continuously if the application genuinely needs it, but do not let production weights mutate continuously without a transaction boundary.

Read the Full White Paper ↗
12 / SELECTED SOURCES

Primary and official references listed in the paper.