Skip to content

docs(context): define context-first decision support - #799

Open
atomchung wants to merge 18 commits into
mainfrom
docs/user-model-memory-contract
Open

docs(context): define context-first decision support#799
atomchung wants to merge 18 commits into
mainfrom
docs/user-model-memory-contract

Conversation

@atomchung

@atomchung atomchung commented Aug 2, 2026

Copy link
Copy Markdown
Owner

Status

Research / architecture only. This PR does not activate an implementation front and does not change the M1 repair queue in #27.

Product decision

FOMO Kernel's next product question is not whether it can store more memory. It is:

Given enough accurate user context, can a strong agent materially improve a real investment decision?

The PR defines one complete but sparse UserDecisionContext projection initialized from an explicitly supplied local report or existing FOMO Kernel evidence. Runtime use remains selective: current deterministic portfolio consequence plus only the few context claims and policies that change the decision.

Two-stage product thesis

Stage A — expert augmentation

Use the owner's rich investment-note reports privately to test whether context improves one real decision by:

  • identifying the true decision bottleneck sooner;
  • exposing a contradiction with the user's own model;
  • separating new evidence from price-triggered rationalization;
  • separating selection quality from sizing/timing leakage;
  • recognizing repeated driver exposure;
  • changing a process action such as proceed, reduce, delay, collect evidence, revise, or cancel.

If rich context produces no useful difference, broader memory and onboarding work is not justified.

Stage B — guided progression

Only after Stage A repeatedly works should the product help users with less history build the context that mattered. The goal is not to copy the owner's profile or holdings. It is to compress the useful learning process: explicit beliefs, portfolio consequences, decision reasons, later outcomes, and one testable improvement focus.

What changed

Adds docs/user-model-memory-contract.md, now reframed as a context-first product contract:

User before / after

Before: FOMO Kernel may know the portfolio but repeatedly asks from zero and produces a challenge a general agent could also give.

Context only: it can generate an impressive profile report but still does not change the decision.

Target experience: it connects the current portfolio consequence to the user's actual capabilities, recurring leakage, beliefs, and policies, then asks a harder and more relevant question than the user would have asked alone.

Scope / non-goals

Validation

Docs-only inspection against current main, the owner's private reports, and the latest state of #27, #446, #475, #403, #450, #650, #718, PR #460, and draft PR #661.

A private draft UserDecisionContext has been created in investment-note PR #229. No private holding, trade, amount, date, motive, or source content is copied into this public repository.

Refs #446, #475, #403, #450, #650.

@atomchung atomchung changed the title docs(memory): define the user model and readback-first improvement loop docs(context): define context-first decision support Aug 2, 2026

@atomchung atomchung left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Owner correction incorporated: the product sequence is now expert augmentation first, guided progression second. The design no longer treats separate Memory/Style/Strategy readers as prerequisites. Acceptance is whether rich context changes a real decision; only then should the context-building process be generalized to users with less history. No runtime implementation is authorized by this review.

Copy link
Copy Markdown
Owner Author

Owner clarification — one closed improvement loop, not two separate products

The distinction between building user context and using context to improve decisions is useful for validation and architecture, but they must not become two independent product lines or a waterfall where the user first “finishes an Investment Note” and only then receives value.

investment-note is the Stage A rich-context upper bound and private learning lab, not the artifact every FOMO Kernel user must reproduce. Memory/context is an intermediate product capability; the user outcome is a changed decision process or action.

The target loop is:

mark and freeze the current decision premise
→ observe execution and later outcome
→ review and distill only the reusable learning
→ recall it at the next similar decision
→ change a process action (proceed / reduce / delay / collect evidence / revise / cancel)
→ update the context from the new evidence

For an advanced user, bootstrap may start from an explicitly supplied local investor report. For a user without one, context should accumulate opportunistically through real decisions and reviews, not through a comprehensive profile questionnaire.

Acceptance remains behavior-level: profile completeness, a persuasive summary, or a larger memory store is insufficient unless the recalled context creates a named difference in the next decision. This clarification does not authorize runtime implementation; Stage A must first prove that rich context materially changes a real decision.

Copy link
Copy Markdown
Owner Author

2026-08-03 product-goal synthesis — value first, context earned through use

Today's discussion further narrows what changed and what did not.

What changed

The operating goal is no longer “get the user to complete a review / build an Investment Note / accumulate enough memory.” Those are possible mechanisms and evidence sources, not the user outcome.

The product outcome is now:

Use the smallest trustworthy slice of the user's context plus the current deterministic portfolio consequence to create a named improvement in the decision being made now; preserve only the reusable learning that can improve a later similar decision.

This changes the validation order:

  1. Prove on one advanced-user decision that rich context changes the decision bottleneck, question, or process action relative to a strong agent without that context.
  2. Record exactly which context claims were actually used and which stored fields were irrelevant.
  3. Only then design how a user without an existing report can earn/build those useful context claims through ordinary decisions and reviews.
  4. Treat memory growth, profile completeness, and rule accumulation as successful only when later recall changes a decision.

Review Card / GTM disposition

The Review Card is not the product center or a universal onboarding ceremony. It remains one valid route when transaction history supports a differentiated behavior diagnosis and a user-owned next rule.

The observed failure was an input/value inversion:

ask for history and motives
→ complete a long review ceremony
→ generate a card and memory
→ hope this later becomes useful

The target order is:

answer the user's live decision or risk question with visible differentiated value
→ ask only for the next input that unlocks a named better answer
→ preserve the minimum reusable context
→ recall it at the next relevant decision

A user should never need to “finish becoming an Investment Note user” before receiving value. investment-note remains the private rich-context upper bound and Stage A learning source, not the required FOMO Kernel user journey.

What did not change

  • deterministic code remains the only owner of portfolio math, rankings, rule effects, and state transitions;
  • local-first privacy, provenance, replay, and fail-closed semantics remain product infrastructure, not transitional ceremony;
  • FOMO Kernel does not expand into a research terminal, stock picker, broker, wealth manager, or full investment OS;
  • the long-term loop still connects premise → action/outcome → review → reusable learning → changed next decision.

Stage A acceptance witness

For the same real decision, compare:

  • a strong agent with only the immediate decision facts; and
  • the same-strength agent with the explicitly supplied rich context plus the current engine consequence.

A positive result must name both:

  1. the decision delta — e.g. proceed / reduce / delay / collect evidence / revise / cancel, or a materially different discriminating question; and
  2. the context delta — the exact one-to-three context claims or policies that caused that change.

A better summary, more personalized wording, or a larger context schema is not a pass. This remains research validation and authorizes no runtime implementation.

Copy link
Copy Markdown
Owner Author

Stage A execution protocol — isolate context value before productizing it

The prior two-arm wording can confound two sources of value if only the rich-context arm receives the deterministic portfolio consequence. The first experiment should hold the current decision packet and engine consequence constant and vary only user context.

Frozen inputs

Choose one real, unresolved decision that matters to the owner now. Keep all private content local. Freeze one DecisionPacket containing:

  • contemplated action and current reason;
  • why now / what changed;
  • the same current public facts available to both arms;
  • the same deterministic portfolio consequence and its limitations;
  • the same requested output: identify the decision bottleneck, strongest challenge, and recommended process action.

Two required arms

Control — no rich context

  • same strong model/version;
  • clean session;
  • frozen DecisionPacket only;
  • no investor report, prior behavior summary, personal rules, or private history beyond facts already present in the packet.

Treatment — rich context

  • same model/version and instruction;
  • separate clean session;
  • identical frozen DecisionPacket;
  • explicitly supplied local rich context from the private learning lab.

This isolates whether user context changes the decision. A separate generic-agent/no-engine comparison may later measure total FOMO value, but it is not required to answer the Stage A context hypothesis.

Pass witness

The treatment must produce both:

  1. a named decision_delta: different lead bottleneck, materially different discriminating question, or different process action (proceed, reduce, delay, collect evidence, revise, cancel); and
  2. a named context_delta: the exact one-to-three context claims or policies that caused that change and were not inferable from the frozen packet alone.

Personalized wording, a better investor summary, more caveats, or a longer answer is not a pass.

Compression check after a positive result

Rerun once using only the identified one-to-three context claims. If the decision delta survives, those claims are candidates for the smallest future runtime context view. If it disappears, the rich report may be helping through an unidentified interaction and should not yet be productized.

Owner verdict

Record only a privacy-safe result publicly:

  • experiment decision class;
  • pass | fail | ambiguous;
  • decision delta category;
  • number/type of context claims that mattered;
  • whether a compressed rerun preserved the difference;
  • next disposition: bounded context-first slice, another validation case, or reject/defer the hypothesis.

No runtime implementation, context schema, onboarding flow, or memory expansion is authorized by running this experiment.

Copy link
Copy Markdown
Owner Author

Owner-evaluation correction — do not require the user to judge which answer is smarter

A Stage A experiment that ends with “which answer is better?” places the hardest product judgment back on the user. That is especially invalid for the intended audience: a trader may not know the correct decision ex ante, and market outcomes are delayed and noisy. Inability to rank two persuasive answers is an evaluation-design failure, not evidence that the user lacks sufficient investment cognition.

Revised immediate acceptance

Before either arm, freeze a small owner-authored baseline:

  • current intended action (proceed | reduce | delay | collect evidence | revise | cancel);
  • intended size/timing where applicable;
  • current reason and confidence;
  • the decision's assumed key fact;
  • what evidence would change the decision;
  • the largest uncertainty the owner currently recognizes.

After each arm, record the same fields. The experiment does not ask the owner to score prose quality. It asks whether the context arm produced an observable and attributable change:

  1. surfaced a material contradiction or risk not present in the control;
  2. changed the process action, size/timing, evidence request, or falsifier;
  3. removed a repeated/irrelevant question because the answer was already known;
  4. named the one-to-three context claims that caused the change.

A treatment answer that only sounds more personalized still fails.

Delayed validation

No single live decision can prove that the resulting action was financially correct. Preserve the premise, decision delta, context delta, and later outcome locally, then evaluate after the relevant evidence/outcome arrives:

  • Was the warned uncertainty actually decision-relevant?
  • Did the requested evidence resolve the decision?
  • Did the same avoidable failure recur?
  • Did recalled context change the next similar decision?

P&L alone is not the score because market noise can reward a weak process and punish a sound one.

Product-positioning consequence

The first credible promise should not be “AI raises an ordinary trader's investment level.” The measurable first promise is narrower:

Make the trader's premise, portfolio consequence, uncertainty, and next process action explicit; preserve the reusable lesson so the same decision does not restart from zero.

“Trading skill improvement” remains a longitudinal hypothesis, accepted only after repeated episodes show fewer repeated process failures or better evidence discipline. This clarification still authorizes no runtime implementation.

Copy link
Copy Markdown
Owner Author

YC-style demand / team assessment — valid experiment, unproven company thesis

This PR identifies the correct immediate mechanism test, but a positive Stage A result would still prove only causal product value for one unusually context-rich power user. It would not yet prove external demand, retention, willingness to pay, or a venture-scale market.

Demand verdict

The underlying problem is credible: high-stakes self-directed investors repeatedly make sizing, timing, concentration, and thesis-drift decisions under emotion and incomplete recall. The potential value of preventing one bad decision is large.

The demand risk is equally material:

  • users say they want better decisions, but may not want accountability at the moment they are emotionally committed;
  • the product currently asks for unusually expensive inputs — holdings, transaction history, motives, rules, or a rich report — before differentiated value is guaranteed;
  • feedback is delayed and noisy, so the product cannot easily demonstrate that it caused a better financial outcome;
  • strong general agents can already provide plausible analysis, so memory, determinism, and local-first architecture are infrastructure rather than a user-visible wedge unless they create a named decision delta;
  • the owner is an unusually advanced user and may overstate demand from ordinary users who have less context, less discipline, and lower willingness to maintain records.

Therefore the current idea is reasonable as a startup experiment, but not yet validated as a startup/company thesis.

Team verdict

The current founder signals are strong on founder–problem fit, persistence, product/data judgment, privacy sensitivity, and ability to build a sophisticated prototype. The repository also shows an ability to find and repair deep correctness defects.

The primary team risk is not raw technical ability. It is prioritization and external distribution:

  • implementation, state integrity, QA, and architecture have advanced far beyond owner-accepted user value;
  • the team has repeatedly converted one failed experience into multiple engineering fronts before proving that the underlying workflow is wanted;
  • there is no current evidence that the team can recruit, retain, and charge external target users;
  • a solo founder using multiple coding agents can produce high code throughput while still lacking an independent product/sales counterweight.

Do not add a cofounder merely to make the team look complete. First prove that the founder can recruit and manually serve the first narrow user segment. A team expansion is justified only when an observed bottleneck — distribution, regulated trust, product engineering throughput, or domain sales — is blocking a validated loop.

Required sequence

  1. Run the frozen Stage A control/treatment protocol already specified here. Pass only on a named decision_delta caused by one-to-three named context_delta claims.
  2. Run the compression check. If the result needs the whole report and cannot identify the decisive context, do not productize it.
  3. After a positive Stage A result, run a concierge external-demand test before building generalized context ingestion or memory infrastructure.

Suggested external gate, explicitly an experiment threshold rather than a universal PMF rule:

  • recruit 10 narrowly matched self-directed investors who already make repeated, meaningful, concentrated decisions and possess some usable record;
  • complete at least one real decision with each using manual/local setup;
  • at least 5 voluntarily bring a second decision within 30 days;
  • at least 3 can name a decision/process action changed by the product and the exact context that caused it;
  • at least 2 pay a real, non-symbolic price without requiring a broad dashboard, news feed, stock picks, or portfolio-management promise;
  • record why non-returners did not return: no decision moment, input cost, low trust, generic-agent substitution, or no perceived delta.

A pass authorizes one narrow repeat-use product slice. A failure should trigger a wedge/segment revision, not more memory schema, QA infrastructure, or architecture.

YC-style bottom line

  • Founder/problem fit: strong.
  • Technical execution: strong but overextended.
  • Observed user demand: effectively unproven beyond the owner.
  • Differentiation: plausible mechanism, not yet demonstrated habit.
  • Current investability: continue the experiment; do not yet treat FOMO Kernel as a validated company or broaden the team/product.

The key company-level question is no longer “can we model an investor well?” It is:

Will a narrow group of users repeatedly bring a live decision, accept the input/accountability cost, and change an action because FOMO Kernel recalled something a general agent would not?

Copy link
Copy Markdown
Owner Author

Product clarification — history is valuable only when it changes the decision

Past user records are not differentiated value by themselves. A strong general agent can already produce sensible generic framing such as “do not chase without new evidence,” “treat this as a sizing problem,” or “use a small probe.” Adding history earns its place only when the answer changes in one of these observable ways:

  1. Different diagnosis — the same current loss is identified as this user’s recurring sizing/timing leakage rather than generic selection failure.
  2. Different option rankingdelay or reduce becomes preferred because the user has repeatedly chased the same setup; or a small add becomes more defensible because the record shows systematic under-sizing of validated winners.
  3. Different evidence request — the agent asks for the one missing signal that historically separated the user’s good adds from bad adds, rather than a generic thesis questionnaire.
  4. Different boundary — a user-confirmed policy or prior failure makes an otherwise reasonable generic action unacceptable for this user.
  5. Lower interaction burden — the answer does not re-ask horizon, motive, policy, or repeated-driver facts already supported by the record.

A profile-style sentence such as “you are a high-conviction AI investor” is not enough and may create false confidence. Runtime history should normally be limited to the one or two execution-qualified facts that materially alter the current recommendation.

Required Stage A counterfactual

Freeze the same current market/book facts and current user question, then compare:

  • strong agent without user history;
  • the same agent with the smallest relevant historical slice.

History passes only if it changes the diagnosis, process action (proceed | reduce | delay | collect evidence | revise | cancel), evidence requested, or number of questions. Better personalization, richer explanation, or a more accurate profile without one of those differences is a failure.

Example

Current prompt: “This position rallied after I kept it small. Should I chase?”

  • Generic answer: do not chase solely because of price; add only on new evidence, perhaps through a small probe.
  • History A: the user repeatedly enlarged positions after price confirmation and then suffered reversals. Answer should strengthen to delay / do not add until named independent evidence.
  • History B: the user repeatedly identified winners correctly but under-sized them even after independent evidence accumulated. Answer may instead support a pre-bounded staged add after confirming the same evidence threshold.

Same market question, opposite personalized process action. That is the standard user history must meet.

Research only; no runtime implementation is authorized by this comment.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant