Skip to content

Latest commit

 

History

843 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Project Theseus

Project Theseus is the experimental implementation of The ASI Stack. Its purpose is to turn the book's highest-leverage mechanisms into causally active software, test them under adequate matched controls, and return honest claim-level evidence to the book.

The active question is:

With the underlying model held fixed, which ASI Stack subsystems improve useful-safe performance on autonomously acquired, machine-verifiable work, why do they help, and under what operating limits?

Theseus is not currently a personal-assistant project and is not yet training its own production model. Autonomous repository work is an experimental substrate. The immediate product is a trustworthy proving system.

Program Selection

Track State Purpose
ASI Stack subsystem proof Active flagship Test one book mechanism at a time with a frozen local model, strong controls, blind evaluation, and source-disjoint qualification
Neural seed Held at step 11,992 Preserve the modular-versus-dense experiment until the subsystem architecture is stable enough for training to answer the right question
Maintenance Service only Protect source binding, evidence custody, storage, replay, and test integrity for the two research tracks

The canonical roadmap defines the re-entry contract. Hardware availability alone cannot resume neural training.

Current Evidence Boundary

  • The ASI Stack experiment identities are pinned to the 84-chapter book at commit 17c6ece80f771d3bce5f89c6b85c99ca9b6c2ea0. Later book work is observed as drift; it does not silently rewrite an opened experiment.
  • The active local controlled variable is mlx-community/Tmax-9B-MLX-8bit at revision 33812d6cf04f88856f25eb828de4f3144a194560.
  • P1 proves that direct and integrated requests traverse the real local model and registered runtime routes. It proves mechanics, not usefulness.
  • Historical P3 solved 1/10 tasks in each route. Its project-selected output ceiling makes it bounded historical evidence rather than current capability evidence.
  • The completed P4 Semantic-IR campaign is INCONCLUSIVE_IMPLEMENTATION: direct solved 3/10, the plan control 1/10, and Semantic IR 0/10, but the Semantic-IR path parsed and lowered only 2/10 against its frozen 8/10 mechanics floor. No broad mechanism conclusion is allowed and D1 remains closed.
  • The replacement role-aware Semantic-IR production owner passes 10/10 deterministic conformance fixtures and one source-disjoint project-authored TMax canary with 2/2 naturally completed production-path artifacts. This closed bounded mechanics repair only. Its prospectively audited fresh-v6 adequacy run then sealed one candidate before a 45,113-token Task 2 prompt produced zero tokens at the 600-second host wall. The exact implementation is terminal INCONCLUSIVE_EXPERIMENT for this TMax/host block; no hidden evaluator ran and no broad Semantic-IR negative is allowed.
  • The sole active claim is now virtual-context-abi.core, selected because the terminal residual is model-visible context materialization and prompt-ingest burden. The frozen source panel binds nine local instrument/control tasks plus a fresh 53-task powered claim panel. Task 26's exact wheel-only uv canary failed closed on sdist-only proxy-tools==0.1.0, scoped to instrument policy. K2.04 subsequently qualified representative parent-only materialization and removed target-derived allowed_effect_paths; K2.05 now has a role-audited target-free 62-row segment plan. All parent stores, static and immutable execution paths, locked closures, and evaluator receipts remain unfinished. No local or Luna claim call is authorized.
  • gpt-5.6-luna at fixed xhigh effort is prospectively defined as a separately denominated OpenAI measurement-only reference. The retained Responses API adapter is offline qualification evidence only and authorizes no VCM calls. Future Luna use requires a demonstrably Codex-subscription-backed route with zero billable API inference; otherwise the arm is omitted.
  • The 57.3M neural shared trunk is preserved at optimizer step 11,992 with append-only checkpoint, optimizer, RNG, and cursor custody. D2 remains unconsumed and capability remains NOT_EVALUATED.

See Project State for the exact current scorecard and Roadmap for the forward sequence.

Scientific Boundaries

The complete operating charter is AGENTS.md. The core rules are:

  1. Test one causally active mechanism at a time. Prompt labels and reports do not count as subsystem execution.
  2. Hold model identity, candidate-visible information, evaluator, effects, and resource opportunities fixed within a comparison.
  3. Keep hidden answers, tests, target-derived metadata, and source identities out of generation and ranking. Recompute integrity independently.
  4. Establish mechanics and implementation adequacy before interpreting a treatment result. An inadequate proxy cannot falsify the full mechanism.
  5. Quality-bearing generation ends on a complete artifact or model EOS. A project-selected generated-token cap cannot manufacture equal quality.
  6. Public benchmarks are calibration-only. Fresh claim-bearing work comes from governed, licensed, source-disjoint repositories.
  7. OpenAI inference is allowed only as a governed teacher or a prospectively sealed measurement-only reference. Luna references must use a demonstrably Codex-subscription-backed route; billable API inference is forbidden. Reference outputs never serve users, train the local model, select tasks, or enter local denominators.
  8. Book support never moves automatically from a Theseus report.
  9. Routine execution must not depend on Corben supplying tasks, labels, approvals, or timing decisions.

Experimental Shape

ASI Stack claim
  -> faithful causal implementation
  -> production-path mechanics and adequacy bench
  -> frozen task/model/evaluator/cost contract
  -> nine-task local adequacy and control screen
  -> powered local VCM vs strongest-control vs flat-ablation comparison
  -> optional separately denominated Luna reference sealed before outcomes
  -> separate conformance/integrity/utility/economics disposition
  -> one fresh D1 qualification for a survivor
  -> backend/model transfer and claim-level book handoff with support unchanged

The machine-readable active state is:

Field Value
Claim virtual-context-abi.core
Phase K3_REAL_WORK_MATCHED_CANARY
State VCM_V4_CHUNKED_ROUTE_PREFLIGHT_GREEN_LARGEST_MATCHED_HOST_PAIR_REQUIRED
Selected task 26
Active attempt vcm_v4_largest_information_matched_flat_and_governed_host_pair_v1
Current wall v4_sub_file_paging_is_call_free_green_and_role_audited_with_six_physically_addressable_vcm_flat_pairs_but_the_new_largest_57626_and_57699_token_matched_pair_has_not_yet_passed_the_non_scoring_host_interlock
Last closed task 26
Next legal action build_prospectively_seal_and_role_audit_one_two_call_non_scoring_v4_host_owner_for_the_new_row4_information_matched_flat_57626_token_prompt_followed_by_the_same_information_governed_vcm_57699_token_prompt_using_the_exact_frozen_TMax_tokenizer_decoder_wrapper_and_600_second_host_interlock_with_complete_artifact_or_model_EOS_and_no_project_selected_quality_token_cap;_stop_on_first_interlock;_never_call_any_consumed_v3_prompt;_authorize_the_nine_task_local_screen_only_if_both_new_receipts_are_GREEN;_keep_Luna_omitted

The VCM instrument is frozen at 62 source-disjoint tasks: nine for local control qualification and 53 for the powered claim campaign. Task 26 was the last bespoke per-task dependency canary; it ended with a role-separated, scoped wheel-only instrument wall rather than a false task or VCM negative. Remaining closures must use one generic manifest-driven owner with shared content-addressed package stores and disposable installed environments. Special handling is limited to Bun, Yarn, TypeScript transpilation, and untrusted Rust risk classes; a storage/host-reserve preflight may stop the instrument before bulk materialization.

The generic owner has replayed the existing npm, pnpm, Cargo, and uv evidence; K2.03 role-separately qualified Bun, Yarn, narrow real-parent TypeScript, and untrusted Rust mechanics; and K2.04 role-separately qualified the production parent-only materializer on four representative rows. The monolithic K2.05 locked batch remains closed by a conservative 40.6 GiB reserve-safe storage deficit. A separate audit has now rederived one target-free 62-row parent manifest and an 8 static / 6 immutable-resolution / 48 locked schedule from the authoritative aligned sources. Panel admission is withheld and no model, Luna, evaluator, repository runner, or dependency execution occurred in the plan. All 62 archive-backed parent stores are now role-separately GREEN across 133,048 regular files and 130,968 UTF-8 pages. Static evaluator Tasks 10, 22, and 27 first qualified; the common target-verifier repair then qualified all eight exact parent-fail/target-pass constructs. All six immutable dependency locks and all six common evaluator environments are now role-separately GREEN. The predecessor matched-verifier campaign qualified Tasks 16, 25, and 56 and scoped Tasks 12, 13, and 35 to host or toolchain walls. One sealed all-or-none replacement campaign has now selected source-disjoint same-panel, same-language successors for exactly those three rows under the original rank rule. Full manifests then correctly invalidated the first Task 13 replacement as explicitly Windows/xloil-bound before dependency execution. A sealed generic host-feasibility successor replaced only that slot with paulomtts/pyjinhx. The authoritative v5 panel and all 124 exact full parent/head closures are now role-audited GREEN. Replacement Tasks 12, 13, and 35 now also have exact locks, offline-replayable environments, and common parent-fail/target-pass evaluators. Together with frozen predecessor Tasks 16, 25, and 56, they form one GREEN, role-audited contiguous six-row K2 instrument. Its parent-only store covers 3,182 regular files and 3,115 UTF-8 pages with 24 independently rederived candidate-visible fields. Controls, descriptive outcomes, cost custody, invalidation classes, complete-artifact-or-EOS completion, and K3 stop rules are frozen with no arbitrary quality token cap. K2 made no model, Luna, teacher, network, or evaluator calls. K3 is still closed until a separate prospective owner binds the exact TMax route, arm packets, blind evaluator, host-operability preflight, one-pass authority, and actual-cost ledger.

K3 begins only after all parent-fail/target-pass receipts, the real parent-only VCM materializer, and the independently recomputed blindness audit freeze under one campaign identity. Its nine-task stage is adequacy—not claim evidence—and selects the strongest eligible local control from flat context, ordinary retrieval, summary/compression, and operable full-parent context without Luna. The 53-task stage compares VCM, that frozen control, and mandatory information-matched flat context. Results are reported separately as L0 conformance, L1 integrity, L2 model use/utility, L3 economics, and later L4 transfer. Luna remains a separate reference denominator and must be rebound to VCM through a proven Codex-subscription-backed, zero-API-spend route before local outcomes or omitted. Evidence volume is not mechanism progress.

Repository Map

Path Purpose
AGENTS.md Durable research, safety, autonomy, and inference rules
roadmap.md Canonical forward execution order
docs/PROJECT_STATE.md Canonical plain-English current state
configs/roadmap_implementation_matrix.json Machine-readable claim, phase, book, and re-entry obligations
configs/project_manifest_registry.json Canonical implementation and route ownership
configs/theseus_external_reference_control.json OpenAI reference-control contract
configs/neural_seed_training_availability.json Neural program and host launch authority
scripts/ Experiment, governance, evaluation, and maintenance owners
tests/ Regression and integrity coverage
reports/ Evidence artifacts; not progress by themselves
runtime/ Private local state and control signals
checkpoints/ Private model, optimizer, RNG, and cursor state

Generated reports, private traces, datasets, credentials, and checkpoints do not belong in a public source release.

Verify The Current Contract

Use Python 3.12+ and a current Rust toolchain. MLX execution requires Apple Silicon and the pinned environment described in the Replication Guide.

python3 scripts/theseus_doc_link_audit.py
python3 scripts/theseus_project_registry.py --gate
python3 scripts/roadmap_implementation_gate.py --gate
python3 scripts/theseus_asi_stack_claim_handoff.py
python3 -m pytest -q tests/test_roadmap_book_sync.py \
  tests/test_roadmap_pretraining_gate.py \
  tests/test_theseus_external_reference_control.py \
  tests/test_neural_seed_training_campaign.py \
  tests/test_neural_seed_autonomous_launch_controller.py \
  tests/test_theseus_asi_stack_claim_handoff.py

Some tests are intentionally artifact-, MLX-, or platform-gated. A skip is not a capability pass. The recenter, registry repair, and source-bound Luna adapter are committed. New active-claim evidence must be rebound once at each coherent source boundary; do not use piecemeal report refreshes to manufacture a green worktree.

For local runtime inspection:

python3 scripts/theseus_assistant_runtime.py --help

Neural Hold

Do not infer training authority from a runnable command. The source-controlled availability policy currently denies launch under HOLD_SUBSYSTEM_PROOF_FIRST, and the runtime yield signal independently asks the segment controller to stop. The checkpoint is preserved, not rejected.

Training can be reconsidered only after the machine-readable Subsystem Architecture Freeze establishes terminal interface dispositions for the architecture-shaping claims, production-equivalent subsystem composition, no unresolved topology/data/route/verifier changes, exact checkpoint rebinding, and autonomous resource and rollback readiness. D2 remains sealed until then.

Documentation Order

  1. Project State
  2. Roadmap
  3. Operating Charter
  4. Glossary
  5. Documentation Index
  6. Top-To-Bottom Architecture
  7. Replication Guide

Honest Walls

  • No ASI Stack mechanism has yet earned source-disjoint D1 qualification.
  • The exact Semantic-IR implementation is frozen for the current TMax/host block after an infrastructure-invalid fresh adequacy run. The broader cognitive-compilation claim remains unresolved, and another current-block Semantic-IR reseal is forbidden.
  • VCM has substantial deterministic mechanics evidence but no prospectively sealed natural-work causal campaign showing that governed context helps the frozen model beyond strong parent-only controls at visible total cost. The v1 candidate protocol's target-derived effect paths must be removed before use.
  • The retained Luna transport is bound to the prior Semantic-IR claim and an API adapter, not VCM or the required Codex-subscription route. No VCM task/evaluator/arm/custody/access contract is prospectively sealed, so no reference call is authorized and later backfill or billable API fallback is forbidden.
  • The neural checkpoint has no capability result, both matched dense controls are incomplete, and D2 is unconsumed.
  • Large generated surfaces, a 97%-used host volume at the review boundary, and broad script/config ownership are real feasibility risks; cleanup must use replay-safe retention owners and forward dependency work must use shared stores rather than per-task cache families.
  • Local security evidence does not qualify LAN or public exposure.
  • The existing public Git history is not an approved release surface.

These are scientific and operational boundaries, not wording problems to hide.

Project Theseus source is Apache-2.0 licensed. Dataset, model, generated artifact, and dependency terms remain governed separately. See LICENSE, Data And Artifacts, and Public Release.

About

Project Theseus public implementation companion for The ASI Stack

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages