Every number your tooling emits is probably correct, and probably not the one you said mattered. Tests passed, runs completed, seconds elapsed: they arrive for free, they go in the report, and the thing you actually set out to change is not among them because nothing produces it automatically. This part measures that distance for this project, in both directions — how much of what Vision was meant to be actually exists today, and how far its own record has drifted from the standard it was given.
This is the last part of The Vision Log, a series about one long-running agent that reads, changes and verifies its own source code. Vision is not an assistant and not an automation framework: it is designed as a digital counterpart to the developer who built it, and autonomy is a means it uses rather than the point of the exercise. Like every part of the series, this one rests on the project’s own engineering log and on documents read directly rather than in summary, and where neither supports a claim the claim is left out. Part 0 explains what the project is and the wish it came from.
Four of the six original goals are live, one is absent, one is half
Grade your own project against what you first said it would be, and the thing you need before anything else is a scale. Vision was originally described as six things. The project’s internal design map places each one, using three verdicts: Live means the running stack goes through that path in production, Unwired means the code exists but nothing calls it except tests, and Absent means it exists in documents only.
| The original goal | Verdict today | Grounds in the engineering log |
|---|---|---|
| Say the purpose, and it picks its own internal tools | Live — one sentence routes to a project, runs it, and walks several gates without intervention | Three entries, 2026-08-17 to 2026-08-19 |
| When it lacks a capability, it builds one | Acquisition Live, creation wired but not yet measured — a wall becomes a search, an install proposal and the developer’s word; a tool it writes and registers for itself has not happened | Three entries, 2026-08-19 to 2026-09-04 |
| Analyse, improve and verify its own structure | Live — and the most mature part of the repository: worktree isolation, gate re-verification, unproven work parked on a held branch | Four entries, 2026-08-25 to 2026-09-03 |
| The model provider is only a brain; identity sits above it | Abstraction Live, memory Live, identity not yet — facts the developer states and findings a run makes are read by every provider in every conversation | Four entries, 2026-08-17 to 2026-09-04 |
| A personal computing environment, an Ubuntu "Vision OS" | Absent — the operating-system adapters contain Windows and nothing else; this one is suspended indefinitely | Two entries, 2026-08-16 and 2026-09-04 |
| Stay aligned with the developer over the long run | Live — their decisions are recorded verbatim with a machine-readable rule, read into every prompt, counted when a judgment cites one, and read back when a rule goes uncited | Three entries, 2026-09-04 |
Four of six are live in some real sense. One is absent. One is half. That is the flattering half of the measurement, and it is the last flattering paragraph in this article. If you ever grade your own project, the useful part starts just after a sentence like that one.
The machines for acting first exist, and almost all are switched off
The machines you build for a system to act first are easy to count and easy to leave switched off. Vision cannot create a tool for itself. The signal is wired: when a model asks three times for a canonical tool that does not exist, that becomes a recorded wall and is handed to the self-fix door (engineering log, 2026-09-04). What has not been measured is the other end — a self-fix run actually writing the tool, and the proof step counting it as applied. The wire has never once fired in over six hundred recorded runs.
Identity above the provider is not built. Rules, the developer’s stated facts, and a run’s findings are all provider-independent memory now (engineering log, 2026-09-04). Vector and graph memory remain two bytes of intent. The abstraction is real; an identity that survives on top of it is not.
Two of five possibility record types never fill. The possibility record accumulates Added, Preserved and Lost from real self-fix judgments; Constrained and Archived, and the expansion runtime behind them, stay Unwired.
Almost every machine for acting first is switched off. The machinery exists: a component that observes friction and turns it into findings, a runtime for self-started initiatives, a self-learning loop, the capability-gap wire above, a record of the developer’s stated decisions, and a record of possibilities a change would open or close. Here is the state of each on 2026-09-11.
| Machine for acting unprompted | State on 2026-09-11 |
|---|---|
| Friction observation | Friction accumulates; the step that turns it into findings last ran 2026-07-20 |
| Friction scanner | Off — it reads a switch nothing in the launcher sets |
| Self-started initiative runtime | Off — same, a switch nothing sets |
| Self-learning loop | Off — the launcher pins it to false at every start |
| Self-task after a blog run | On, and the only one |
| Capability-gap wire | Live, never once fired in over six hundred recorded runs |
| Developer decisions and the possibility record | Live and used daily |
The single path that is switched on fired twice in two days, and produced nothing both times. The first self-task reached a run and came back with a provider quota refusal. The second created a conversation and then never delivered the body of the task into it, so a conversation with zero messages is all that exists of it. Neither failure told anyone anything.
Self-repair is the exception, and it is a real one: project-scope walls open a repair door on their own, a run walks through it, and fixes land without anyone asking. But a wall is something a run collides with. It is reaction, not initiative, and it is the same shape as the problem that retired the system before this one, wearing better clothes. When you count what your own system does unasked, count only what it began with nothing arriving first.
The charter names eight measures, and this series had said five
The measures you wrote down at the start are not necessarily the measures you are using now. So much for what is built: the harder question is what any of it should be judged by, and this project wrote that down before it started building.
The founding document, version 0.2, was ratified on 2026-07-15. Its closing section is headed Success Metric, and it is two lists: what a digital counterpart is evaluated by, and what it is not.
| Evaluated by | Not by |
|---|---|
| Future Possibility | Task Count |
| Growth for the developer | Runtime Count |
| Shared Growth | Model Count |
| Capability Growth | |
| Knowledge Growth | |
| Contribution | |
| Continuity | |
| Trust |
The first thing to report is a correction to this series, and it is the reason to open a document rather than accept a summary of it. Earlier parts announced this subject as five measures: growth for the developer, shared growth, capability growth, continuity and trust. That was wrong, and those parts have since been corrected to match this one. The document names eight. The three the summary dropped are future possibility, knowledge growth and contribution, and one of those three turns out to be the only measure on the whole list that this project takes a reading on today. A summary that loses three items out of eight will eventually lose the item that mattered.
The second thing is a single word. An earlier section of the same document defines growth measurement as evaluating six deltas: capability, knowledge, possibility, performance, trust and autonomy. Delta, not total. A growth measure is a difference between two readings, so an instrument for one has to take the reading twice and subtract. That requirement on its own disqualifies almost every number this log has published, because a count taken once is not a difference.
The third thing is that the definitions are narrower than the names make them sound. Read the definition before you adopt somebody else’s measure, never the word on the label.
- Growth for the developer. The document’s stated objective for long-term collaboration is one sentence: decrease their workload while increasing both their capability and Vision’s. It is about a person’s hours and a person’s reach, not the system’s throughput.
- Shared Growth. Their growth implies Vision’s and Vision’s implies theirs; the relationship is written as symbiotic, under an invariant that Vision never attempts to replace them and continuously amplifies them.
- Trust. It evolves through correctness, transparency, consistency, reliability, learning and recovery, and it decreases through hidden decisions, unsupported claims, repeated failures, unexplained behaviour and policy violations. The document adds a fixed point: trust never becomes complete, it is continuously earned.
- Continuity. The document equates identity with persistent continuity and writes the chain out: identity, then experience, then reflection, then learning, then the next identity. It means a self surviving its own changes, not a process surviving a restart, so ask of your own system which of the two it is actually keeping.
Not one of 1,417 recorded numbers is denominated in any of the eight
Count the numbers your own project records and sort them by what their unit is. The engineering log now holds 531 entries across 28 days of work, from 2026-08-06 to 2026-09-11. For this article every quantity in it that carries an identifiable unit was extracted and sorted by what the unit is. There are 1,417 of them.
| What the number counts | How many such numbers |
|---|---|
| Elapsed time, from milliseconds to days | 410 |
| Discrete items and cases | 410 |
| Tests and assertions passed | 125 |
| Published articles | 115 |
| Files | 78 |
| Failures, errors and warnings | 70 |
| Size on disk | 67 |
| Occurrences of an event | 58 |
| Tests | 54 |
| Runs | 30 |
Grouped more coarsely: 940 of the 1,417 count things, 410 measure elapsed time, and 67 measure size on disk. Not one is denominated in any of the eight measures the charter names. Task Count and Runtime Count, the first two entries on the forbidden list, describe the entire vocabulary.
The other parts of this series inherit that vocabulary, because they were written out of that log. Listing every distinct thing they measured gives twenty-six subjects, and they sort like this.
| What a part of this series measured | How many | Which list it belongs to |
|---|---|---|
| Work items counted: log entries, held topics, source documents, extracted propositions, swept files, recorded runs, objections, and the tasks and steps left behind by two abandoned systems | 17 | Task Count |
| Runtime counted: seconds a launcher call hung, minutes of downtime, process identifiers named in a cleanup error | 3 | Runtime Count |
| Configuration states: which loops are switched off, how many grounds a fix must satisfy, which machines for acting unprompted are enabled | 3 | Neither. A setting, not a measurement |
| Capability inventories: where each of six original goals stands, how many capabilities were carried forward from the earlier systems | 2 | Names capability, but takes one reading and never a second |
| A note that two of five possibility record types never fill | 1 | A report on an instrument, not a reading from it |
Twenty of twenty-six are the forbidden list by name. Not one of the twenty-six is a reading on any of the eight. Count your own dashboard the same way and you will probably find the same split: plenty of numbers, none of them denominated in the thing you said mattered.
Two of the eight words have large implementations that mean something else
You can hold a page of correct numbers and still have no instrument for the thing you said mattered. The interesting question is not whether those numbers are wrong. They are correct numbers about real things. The question is whether an instrument exists for the things the charter actually asks about. Each of the eight was checked against the repository as it stood on 2026-09-11.
| What the charter names | Is there anything that measures it |
|---|---|
| Future Possibility | Yes. Since 2026-09-05 every self-repair commit is read for the options it added, kept and removed, and each judgment leaves one line in a possibility record. 77 readings so far |
| Growth for the developer | No. Nothing measures their workload or their capability. The word workload occurs in the source only about compute |
| Shared Growth | No. The nearest record is a count of times a judgment cited one of the standing decisions: 97 citations against 14 recorded decisions. That measures whether a rule was read, not whether either party grew |
| Capability Growth | No delta. There is an inventory of what is live, half-built or absent, refreshed by hand. Nothing takes that reading twice and subtracts |
| Knowledge Growth | No. There are knowledge stores and their sizes are countable, which is the trap the project’s own risk register named before any of this was built: data volume mistaken for intelligence, with behavioural and outcome metrics required before a growth claim |
| Contribution | No. The word occurs in the source only as unrelated field names |
| Continuity | Not this continuity. Code carrying that name measures failover, checkpointing and migration, a process surviving an interruption. The charter’s continuity is an identity surviving its own learning |
| Trust | Not this trust. Code carrying that name scores how far an observation source can be believed. The charter’s trust is the developer’s, and it has no reading at all |
The last two rows are the ones worth sitting with. Search the repository for the charter’s words and two of them come back with substantial implementations attached, hundreds of lines each, well tested. Both are homonyms. A trust score for a data source and a continuity guarantee for a service are useful things to own, and neither is the thing the charter is asking for.
That failure mode deserves a name, because it is quieter than an absence. A word with no implementation is visibly missing. A word with the wrong implementation looks handled, and it goes on looking handled for exactly as long as nobody reads the definition next to the code. Open two of your own words against their code this week, and start with the one that looks handled.
The one instrument, and the series that never quoted it
Future possibility is the exception, and it is a real one. Between 2026-09-05 and 2026-09-11 the possibility record accumulated 77 entries, one per judged self-repair: 16 options added, 149 preserved, 0 lost — which is a live reading on one of the charter’s own measures, and not one article in this series had quoted it before this sentence.
Zero lost is not luck. A repair that removes an option, whether a module, a route, a tool, a fallback provider or a recovery path, is held out of the main line whatever else it proved, until it is released deliberately. A lost option stops the change rather than being recorded against it, so the zero in that column is the gate working rather than the absence of the problem.
That omission is the uncomfortable part, and it is worth more than the reading. Four parts about what this project measures, a reading on one of the eight taken 77 times over a week, and the one measurement that satisfies its own charter went unmentioned in all of them. It was not hidden. It is small and recent, the counts are large and old, and reaching for the large old numbers requires no decision from anyone, which is exactly how a project ends up measuring diligently against the wrong list. When one of your measures is small, recent and awkward, quote that one: the large old numbers will get quoted for you.
Fifty-one days passed before the standard was held against the record
Nothing will schedule the hour in which you hold your project against its own standard. The charter was ratified on 2026-07-15. The engineering log’s first dated entry is 2026-08-06. The first entry to put the two lists side by side and state that the log measures almost entirely the forbidden one is dated 2026-09-04.
Fifty-one days between writing the standard and comparing anything to it, with three weeks of daily, numbered logging inside that gap.
Nothing was concealed during those fifty-one days. Every forbidden-list number was in the open, dated and attributed. The comparison needed no new data and no new tooling. It needed someone to open the standard and the record at the same time, which turns out to be a different act from producing either one, and a much rarer one. Put that hour in your own calendar, because a working project will never schedule it for you.
Four mechanisms landed on the last day, and every reading is zero
A new capability reaches you as code long before it reaches you as behaviour, and only the second is worth anything. Four mechanisms landed on 2026-09-11, aimed at the gap between a tool that waits and a colleague. They are reported here with their readings, and the readings are all zero.
| What landed on 2026-09-11 | What it begins to measure | Readings so far |
|---|---|---|
| A line a run may write to object to its own instructions, and a second to recommend a course nobody asked for | Whether Vision is worth listening to. Each one is left open and later marked right, wrong or moot, and the ratio is right over judged | 0 |
| A board of what the project is currently carrying, assembled from the stores that own each part and read into every run’s instructions | Whether work that gets picked up gets finished | 0 |
| A review of the project’s own operating record that may open one task on the single axis that got worse against its own five-day median | Whether the system moves before it is pushed | 0 |
| Durable delivery for a self-started task, so the body of one cannot die in memory while it waits for the machine to go idle | Nothing. It is a repair, made after two self-started tasks died silently in two days | Not applicable |
The baseline they were built against was measured on the morning of 2026-09-11 across the 614 runs recorded since 2026-08-17. The markers the prompt explicitly asks a run to emit are plentiful: over a hundred runs reported something learned, and eighty-six recorded a wall they could not pass. Those measure compliance. The unprompted categories are these: four runs objected to their own instructions, eight offered a recommendation nobody asked for, and three did work outside the scope they were handed. Fifteen in total, 2.4 percent. Over the most recent week of that window, one objection and no recommendations at all.
One detail is deliberate and worth reporting. The review that can open a task on a worsening axis was written by the project, but the switch that enables it was not: it sits in machine-scope configuration, outside what the system may change on its own. It is on, and a person turned it on. A system that can widen its own initiative without anyone’s word has widened something other than initiative.
So the honest statement of where this stands is not that the chair is filled. It is this: the wish was for a colleague rather than a tool that waits, Vision is what has been built toward that, the parts that let it act first are largely present, and almost all of them are switched off or have never fired. Four mechanisms aimed at exactly that gap were written on the day this series ended, and not one of them has been exercised. None of it is even running yet — a change here is not live until a person restarts the process. Read any claim of a new capability, in this series or anywhere else, as a claim about code that exists rather than about behaviour somebody has watched.
What changes for you
You can adopt the practice this series argues for in an afternoon, and it is easy to skip, which is why it took fifty-one days here.
Write the standard before the instrumentation, and date it: a standard written afterwards will describe the instruments you already have, and the only reason the comparison above is possible is that this one was ratified three weeks before the first log entry. Then put the comparison itself on a schedule, because producing a record and auditing that record against a standard are separate activities and only one of them happens by itself. While you are there, check your definitions against your identifiers — two of the eight words here had large, well-tested implementations that meant something else entirely, and a word with the wrong code looks finished in a way a missing word never does. Remember that growth is a delta, so build the second reading first: an instrument that cannot take the same measurement twice and subtract is measuring a state, and a state cannot tell you whether anything grew. And suspect the metric that arrives for free, because test counts, run counts and elapsed seconds are produced by the tooling whether or not anyone wanted them, while anything about the person you are building for has to be built on purpose: across the 614 runs counted above, the markers this project’s own instructions ask a run to emit are everywhere, and the categories nobody asked for are 2.4 percent of the total.
What holds the five parts of this series together is not the machinery. Each one is an account of being wrong in a way that was legible only afterwards, and legible only because the wrong version was kept: a system that passed every test it had and was abandoned anyway, two safety checks that behaved exactly as designed and produced seventy-eight minutes of silent downtime between them, four days of explanations that measured flat before the real defect turned up four steps upstream. This part is the same shape one level up, a project measuring diligently, accurately and in detail against a list it had itself ruled out in writing.
None of that is a story about carelessness. Each is a story about a check nobody had a reason to run until it was overdue, and the single practice underneath all of them is keeping the wrong version next to the right one, dated, so the distance between them can be measured later by someone who was not there. That is the whole method, it is the only reason any number in these five articles could be checked, and it is the part you can take away once every specific here has gone stale.
Where this goes next
This is the end of The Vision Log: five parts, written over two days, from a record kept by the system they describe. The next series moves from the agent to what it builds — four game projects and the 303 recorded runs behind them, which is a far messier record than this one and a more honest test of whether any of this machinery produces something a person would want to play. It opens with what that record cannot tell you, and is here.
Part 0, on what Vision is and the wish it was built from, is here. Part 1, on the two systems built and abandoned before this one, is here. Part 2, on what breaks when an agent edits its own repository, is here. Part 3, on four wrong diagnoses before the real one, is here.
