A plan that never looked at the scoreboard
Agentlas rooms can pursue one Goal for days or weeks. The host keeps the plan as a tree: a mission with key results, a few strategies under it, and tactics under each strategy. One of our own rooms runs a Goal to grow a small Threads account to 1,000 followers in a month.
The tree was drafted well on day one. After that, only the bottom of it moved. In two weeks the room made 23 changes to its plan, and every one of them was a tactic. Thirteen of those changes were refused because the strategies already held the maximum of twelve open tactics, so even the attempts to change direction bounced. The room ran nine strategy reviews, and all nine recorded the same thing: no sensor. Meanwhile the account grew by about one follower a day against roughly 45 a day needed. The four strategies stayed as they were first written.
Two causes
Strategy reviews were scheduled by time, not by what the numbers did. A review that fires every few days has nothing to say when nothing in it has changed.
The numbers were not on the host. The agent read the follower count on screen and wrote it into the conversation, but nothing tied that sentence to the key result it measured. A review had no series to look at and no way to tell a bad day from a stalled month.
What Agent Strategy adds
Agent Strategy puts a metrics layer on top of the existing plan tree. It uses only Goal-scoped records and is separate from the memory system.
- A KPI time series per key result. A Goal turn reads the real source and reports the value with the time it was read, where it was read and what showed it. The host validates the report before storing it.
- A pace state. The host compares the observed pace with the pace the target requires and classifies the key result as on pace, ahead, behind, stalled, declining or breakout.
- Transition triggers. A re-plan is requested only when the state changes and the change is confirmed, never on a timer.
- A structured re-plan. The next turn receives the evidence and answers in a fixed format: keep, replace or add a strategy, or report to the owner.
- Receipts. Every applied change records which samples justified it, what it retired and what it added.
Architecture
The host owns everything that can be checked: the samples, the state, the gates and the final application. The agent owns the judgment: which strategy to keep and what to try next.
Deciding that something changed
Follower counts are small integers read once or twice a day, so a naive rule would flip between states on a single sample. Four rules keep the state honest.
Two windows. The long window is 72 hours by default and widens, up to 14 days, when samples are sparse. The short window is a third of it, at least 24 hours. A stall in the long window that the short window contradicts is withheld. This is the multiwindow pattern SRE teams use for burn-rate alerts.
Hysteresis. Each state has a wider band to stay in than to enter, as in the table below, so a value sitting on a boundary cannot flip it back and forth.
Confirmation. A state changes only after two consecutive evaluations on different sets of samples agree.
Enough data. Fewer than three samples, or samples covering less than half the window, give no data. No data is reported as a measurement gap, never as a failing strategy.
| State | Enters when observed ÷ required pace is | Stays while it is |
|---|---|---|
| Stalled | below 0.3 | below 0.5 |
| Behind | 0.3 to 0.8 | 0.2 to 0.9 |
| On pace | 0.8 to 1.25 | 0.6 to 1.4 |
| Ahead | 1.25 or more | 1.0 or more |
| Breakout | short-window pace at least twice the previous long-window pace and at least 0.8 of the required pace | while it stays at 1.3 times the previous pace or more |
Limits that keep it from thrashing
Re-planning too often is as harmful as never re-planning. A new strategy needs time to show anything, so the gates are deliberately conservative.
- At most one trigger a day and three applied strategy changes in seven days.
- A cooldown of at least 48 hours after a trigger, doubled when the same state fires again.
- A trigger needs samples from at least two different runs, and it expires after 72 hours if no turn acts on it.
- A re-plan may make at most two structural changes and may not touch a strategy that is still inside its protection window.
- The host checks that every sample the re-plan cites exists. A low-evidence re-plan cannot claim high confidence.
- When the required pace is ten times the observed pace or more, the host does not ask for another pivot. It reports the current value, the required pace and the observed pace to the owner and asks whether to adjust the target.
- If the owner pauses the Goal, nothing fires.
A strategy change is one transaction
The old route changed the plan one operation at a time, and the twelve-tactic cap was checked after each operation. That is how the room above kept bouncing off the cap. A strategy change is now applied as a whole: retiring a strategy closes its open tactics, and adding a strategy opens its new tactics, in the same transaction. If any part fails, nothing is applied.
Replaying the room's own history
Before turning it on we replayed the follower counts this room had written into its own conversation, from September 27 to October 6, through the new rules. The state became stalled by October 1 and the host fired exactly one trigger: the owner report, because about 45 followers a day were needed and about one a day was observed. The original system kept all four strategies unchanged for the following twelve days.
Those readings were taken from conversation text, not from a host ledger, which is exactly the gap this work closes. The replay is part of the release contract test, together with 19 other checks on states, gates, validation and atomic application.
First production measurements
Agent Strategy shipped in Agentlas Desktop 1.2.84 and is on by default. On October 11 at 08:40 KST the same Threads room ran its first scheduled turn with it. When the turn rendered the plan, the host created the key result's KPI: a target of 1,000 followers by October 28, a measurement every 24 hours, and three samples before any state is assigned.
Two minutes later the turn read the profile in the shared browser and reported 31 followers, with the evidence that the browser snapshot displayed 31 followers. The host validated the report and stored it as the first sample, tied to the run that read it. Until this release, that number lived only in the conversation text.
The state is no data, by design. One sample says nothing about pace. With one reading a day, the earliest classified state is October 13 and the earliest confirmed trigger October 14. We will add the first live state change and re-plan to this note when they happen, with their receipts.
Where the design comes from
The thresholds come from these sources and from one account's data. They live in a single table so they can be tuned as more Goals run.
- Change-detection bandits such as CUSUM-UCB, which restart a policy only after a detected change instead of on a schedule.
- SRE multiwindow alerting, which pages only when a long and a short window both cross the threshold.
- Plan-and-Act and ADaPT, which update a plan from execution state and decompose only the part that failed.
- Growth practice: a North Star metric broken into inputs, and ICE scores for choosing the next experiment.
Limits
The values are read by the agent from the screen. Evidence is required and sudden jumps are held as suspect until confirmed, but this is not as trustworthy as a value from an official API or a collector.
An account-level number cannot say which strategy moved it when several run at once. The re-plan is left to the agent's judgment, and attribution experiments are future work.
One account cannot run a controlled comparison. Any before-and-after improvement we report will be a comparison, not proof of cause.