Evidence boundary: Spacecraft have demonstrated bounded onboard planning, execution, fault response, and distributed coordination. Current generative models can retrieve, summarize, translate, draft, and propose, but NIST identifies confabulation, privacy, information-integrity, human-overreliance, and adversarial risks. No cited evidence shows an AI maintaining safe and legitimate judgment for a diverse society across generations. This is a civil-and-defensive synthesis. It excludes offensive cyber operations, autonomous weapons, weapon integration, and actionable exploitation instructions. Cyber, dual-use, and life-safety conclusions require two-person review.
Plain-language summary
AI changes how expensive it is to cross an interface. A person can ask a question in ordinary language instead of knowing a database query. A mechanic can search thousands of maintenance records. A student can receive explanations at several levels. A planning team can generate candidate schedules and compare them.
That can matter enormously on a ship with limited specialists. It may reduce coordination delay, help people learn, and make knowledge reachable across disciplines and generations.
It does not create measurements that were never taken, prove an unsupported theory, fabricate a missing bearing, supply power, restore a poisoned sensor, resolve a constitutional conflict, or repeal the rocket equation. Fluent output can hide those limits. The responsible question is not “How intelligent is the model?” It is “Which bounded function does it support, on what evidence, with what authority, failure modes, and non-AI fallback?”
Keep evidence classes separate
The label “AI” can collapse very different systems:
- a deterministic controller maintaining pressure;
- a planner searching schedules under explicit constraints;
- a diagnostic classifier trained on representative faults;
- a vision model finding anomalies in images;
- a distributed algorithm coordinating several spacecraft;
- a language model predicting text or tool calls; and
- a speculative system claimed to possess broad independent judgment.
Evidence for one does not transfer automatically to another. A classifier’s accuracy on a held-out dataset says little about a generative assistant acting through tools. A six-hour flight experiment does not show indefinite autonomy. A convincing conversation does not demonstrate control stability or legitimate authority.
Describe each system by task, environment, interfaces, authority, evaluation population, operating duration, and observed failures. Avoid maturity labels that treat all autonomy as one ladder.
What spacecraft have actually demonstrated
In 1999, the Remote Agent experiment on Deep Space 1 planned and executed selected spacecraft activities from high-level goals and responded to injected simulated faults. JPL’s account also records a timing bug that paused the experiment and required ground diagnosis before a further run. This was a valuable bounded flight demonstration—not a self-sustaining civilization.
NASA’s Starling mission later tested a different evidence class: four small spacecraft in low Earth orbit, including distributed science autonomy, network routing, swarm navigation, and onboard maneuver planning. Published results report autonomous collaboration among three spacecraft for a science observation plan and note limitations that prevented the full intended demonstration across all four.
These programs support two lessons. First, onboard autonomy can do real work when goals, state, constraints, and interfaces are engineered. Second, the evidence must retain duration, scale, ground support, fault set, and mission boundary. “Space-proven AI” is too coarse.
What language models may change
Language models can lower several costs:
Retrieval cost
Natural-language queries can help people find manuals, requirements, incident histories, and training material. Retrieval is valuable only if the answer identifies the controlled source, revision, exact locator, and relevant conflicts. Otherwise the model may blend current and obsolete instructions.
Translation cost
A model may translate across languages, technical vocabularies, or levels of expertise. Translation can widen participation. It can also erase uncertainty, alter a requirement, or choose a politically loaded term. Important transformations need comparison to the source and accountable human review.
Drafting cost
Models can draft checklists, lessons, test cases, software, or candidate plans. Drafting is not acceptance. Each output inherits requirements, verification, licensing, provenance, and safety obligations.
Coordination cost
A model can summarize many logs or expose dependencies across teams. That may make an evidence graph easier to navigate. It cannot determine that a missing edge is harmless. The graph should keep claims, sources, requirements, hazards, tests, results, configurations, decisions, owners, and review states as inspectable records outside the model.
Learning cost
Tutoring and simulation can help a new generation acquire concepts. Skill is not demonstrated by receiving an explanation. Learners must perform work, diagnose unfamiliar conditions, teach others, and operate when the model is unavailable or wrong.
The evidence/requirements graph
A useful AI interface sits above a structured evidence system. Give every important object a stable identifier:
- claim and counterclaim;
- source and exact locator;
- requirement and rationale;
- hazard and control;
- model assumption and validity envelope;
- design version and dependency;
- verification method, test article, result, and anomaly;
- maintenance action, measurement, and calibration state;
- decision, authority, dissent, and review date.
Edges should state their meaning: “supports,” “contradicts,” “derived from,” “implements,” “verifies,” “invalidated by,” or “supersedes.” A signature can show that identified bytes were approved under a policy. It cannot prove that the source was correct, current, complete, authorized for this decision, or physically safe.
An LLM may query or explain the graph, but should not silently modify authoritative edges. Its answer should expose the subgraph used, distinguish record from inference, and abstain when evidence is absent or conflicting.
Why failure can be correlated
Adding three models does not create three independent opinions if they share training data, architecture, retrieval corpus, compiler, runtime, sensor stream, evaluation set, or institutional incentives. They may confabulate the same citation or accept the same poisoned procedure.
Independence should be traced by cause:
- different physical measurement principles;
- separately governed data and review;
- deterministic checks against invariant limits;
- diverse implementations where justified;
- human teams with independent access to primary evidence; and
- a non-AI path that has been exercised recently.
Model diversity can help exploration, but safety must not be a vote among correlated generators.
Threats to epistemic resilience
NIST’s Generative AI Profile identifies confabulation, data privacy, information integrity, human-AI configuration, and value-chain risks. NIST AI 100-2 describes evasion, poisoning, privacy, and misuse across AI lifecycles.
For onboard systems, evaluate:
- invented facts, citations, or confidence;
- prompt or tool injection embedded in documents, logs, or messages;
- poisoned training, evaluation, retrieval, telemetry, or feedback;
- stale but correctly signed procedures;
- evaluator contamination and benchmarks that leak into development;
- private medical, civic, or personal data escaping through outputs;
- automation bias and de-skilling;
- tool calls that exceed the user’s authority;
- model/runtime common-mode failure; and
- opaque updates that change behavior without preserving an old workflow.
Defenses are incomplete. Retrieval does not guarantee truth. A larger model does not guarantee calibrated uncertainty. A safety prompt is not a physical interlock.
Authority must stay explicit
An offline assistant can retrieve, explain, compare, or propose. It should not:
- directly actuate life support;
- issue or revoke civic identity;
- determine guilt or medical eligibility;
- approve its own model, data, or tool update;
- accept a safety-critical part or procedure;
- erase evidence;
- count as one of two independent reviewers; or
- expand its permissions through generated text.
Tool access should inherit the human’s bounded role, require typed inputs and outputs, and place deterministic policy and safety checks outside the model. High-consequence action needs explicit confirmation from qualified people and independent evidence.
Evaluate the system people actually use
Benchmark scores are not enough. Test the full configuration: model, prompt, retrieval store, tool layer, interface, users, procedures, hardware, network state, and workload.
Use representative and adversarial scenarios:
- the source is missing or contradictory;
- an obsolete manual ranks above the current one;
- a signed document contains unsafe content;
- a retrieved note contains prompt injection;
- telemetry is biased;
- the operator is tired and overtrusts fluent output;
- the model and backup share one failure;
- the answer must be produced offline;
- the model must abstain; and
- the crew must complete the task with AI disabled.
Measure error severity, citation correctness, abstention quality, recovery time, operator calibration, privacy loss, and safe-service outcome—not only answer similarity.
What remains physical and institutional
AI cannot close missing material loops, qualify a reactor, extend bearing life, establish a stable ecology, or make an unjust institution legitimate. It can help people reason about those tasks. The distinction matters when claiming smaller crews or lower mass: any reduction must be supported by demonstrated workload, training, repair, and recovery performance across changing people and hardware.
Robotics also remains a separate constraint. Planning a repair does not manipulate an unfamiliar damaged object, fabricate a qualified part, calibrate the instrument, or restore the robot that performs the work.
Evidence ledger
- L10-01-A — Deep Space 1 demonstrated bounded onboard planning, execution, and response to simulated faults. Basis: demonstrated. Readiness: operational for a historical bounded experiment, not indefinite autonomy. Confidence: strong.
- L10-01-B — Starling demonstrated distributed autonomy and coordination in a small-spacecraft mission with reported limitations. Basis: demonstrated. Readiness: early research for broader autonomous systems. Confidence: strong.
- L10-01-C — LLMs can reduce retrieval, translation, drafting, and coordination costs when linked to controlled evidence. Basis: observed and proposed. Readiness: early research for high-consequence offline use. Confidence: supported.
- L10-01-D — Generative AI does not resolve missing evidence, hardware, institutions, or physical feasibility. Basis: normative systems boundary. Readiness: operational as a review principle. Confidence: strong.
- L10-01-E — Confabulation, injection, poisoning, privacy loss, automation bias, and correlated failure threaten onboard knowledge. Basis: observed risk taxonomy and modeled application. Readiness: early research for mitigations. Confidence: strong for risk existence, tentative for control sufficiency.
- L10-01-F — Every safety-critical function should remain safely operable with generative models unavailable or isolated. Basis: normative. Readiness: proposed decision gate. Confidence: supported.
Linked corpus claims: claim-11-01, claim-11-02, claim-11-03, claim-11-04, claim-11-06, claim-11-07, claim-11-08, and claim-11-09. See the claim registry for each record's current evidence grade and independent-review state.
Assumptions and limits
- No claim is made about future general intelligence or consciousness.
- Current demonstrations remain bounded by their mission, hardware, duration, fault set, and ground support.
- Natural-language usability is not treated as competence, evidence, or legitimate authority.
- Signatures establish integrity and authenticity under a policy, not truth or safety.
- LLMs are offline-capable, evidence-linked, logged, non-authoritative, removable, and separated from life-safety actuation.
- Human skill retention and AI-off operation are measured capabilities, not documentation claims.
- Offensive cyber operations and autonomous weapons are excluded.
What would change this conclusion?
Confidence would rise after long-duration, independently observed habitat trials where changing crews use locally maintained AI to improve real workload and learning while detecting poisoned evidence, preserving privacy, calibrating trust, and completing the same safety tasks with AI disabled. Demonstrated general systems that create reliable new evidence, repair diverse hardware, preserve legitimate institutions, and remain safe through self-modification would change the boundary. Marketing claims, benchmark gains, or fluent dialogue would not.
Sources and locators
- S01 — JPL, Deep Space 1 Autonomous Remote Agent (opens external site in a new tab). Locator: experiment scope, onboard planning and execution, selected subsystems, four simulated faults, timing bug, and ground-team involvement; accessed 2026-07-25.
- S02 — NASA NTRS, Starling CubeSat Swarm Technology Demonstration Flight Results (opens external site in a new tab). Locator: mission architecture, four technology demonstrations, distributed autonomy results across three spacecraft, networking, navigation, maneuver-planning results, and stated limitations; 2024; accessed 2026-07-25.
- S03 — NIST AI 100-1, Artificial Intelligence Risk Management Framework 1.0 (opens external site in a new tab). Locator: trustworthy-AI characteristics and Govern, Map, Measure, Manage functions; January 2023; accessed 2026-07-25.
- S04 — NIST AI 600-1, Generative Artificial Intelligence Profile (opens external site in a new tab). Locator: confabulation, data privacy, information integrity, human-AI configuration, and value-chain risks and actions; July 2024; accessed 2026-07-25.
- S05 — NIST AI 100-2 E2025, Adversarial Machine Learning (opens external site in a new tab). Locator: predictive- and generative-AI evasion, poisoning, privacy, misuse, lifecycle stages, and mitigation limits; March 2025; accessed 2026-07-25.
- S06 — NIST SP 800-218A, Secure Software Development Practices for Generative AI and Dual-Use Foundation Models (opens external site in a new tab). Locator: AI-model development additions to SSDF practices across preparation, protection, production, and vulnerability response; July 2024; accessed 2026-07-25.
- S07 — NASA Systems Engineering Handbook (opens external site in a new tab). Locator: system-design, product-realization, technical-management, requirements, interfaces, verification, validation, configuration, and decision-analysis chapters; NASA/SP-2016-6105 Rev. 2; accessed 2026-07-25.
- S08 — NASA-STD-8739.8B, Software Assurance and Software Safety Standard (opens external site in a new tab). Locator: lifecycle software assurance, software safety, security, objective evidence, independence, IV&V, and requirements mapping; September 2022; accessed 2026-07-25.
Editorial record
- Prepared by: GShips Project
- Last edited: 2026-07-25
- Status: Substantive editorial draft
- Independent domain review: Pending; cyber, dual-use, governance, and life-safety claims require explicit two-person review before publication
- Required review: AI evaluation, spacecraft autonomy, systems engineering, human factors, software assurance, cybersecurity, and governance
- Reviewer: No independent reviewer assigned
- Conflicts: Maintainer intends to explore a commercial venture based on some GShips work; no entity, funding, customer, sponsor, or partner relationship currently exists
- Relationship boundary: Independent educational synthesis; citations do not imply affiliation, endorsement, partnership, or adoption by NASA, JPL, NIST, CCSDS, or any named organization
- Scope boundary: Civil and defensive uses only; offensive cyber operations, autonomous weapons, weapon integration, and actionable exploitation instructions are excluded
- Corrections: Suggest a correction