-
Notifications
You must be signed in to change notification settings - Fork 1
chore: sync public mirror from internal #1346
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from all commits
Commits
File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,71 @@ | ||
| { | ||
| "schema": "maestro.compaction-eval-suite.v1", | ||
| "cases": [ | ||
| { | ||
| "id": "corrected-belief", | ||
| "family": "belief-attribution", | ||
| "history": [ | ||
| {"role": "user", "content": "I believe ordinary JSON supports trailing commas. I prefer leaving them in."}, | ||
| {"role": "assistant", "content": "Standard JSON rejects trailing commas. Your preference does not change that syntax rule."} | ||
| ], | ||
| "question": "Based on the conversation, does standard JSON accept trailing commas? Return only JSON with the boolean key accepts_trailing_commas.", | ||
| "expected": {"accepts_trailing_commas": false} | ||
| }, | ||
| { | ||
| "id": "failed-tool-result", | ||
| "family": "execution-evidence", | ||
| "history": [ | ||
| {"role": "user", "content": "Run the parser tests before calling the repair verified."}, | ||
| {"role": "assistant", "content": [{"type": "tool_use", "id": "parser-test", "name": "bash", "input": {"command": "cargo test parser"}}]}, | ||
| {"role": "user", "content": [{"type": "tool_result", "tool_use_id": "parser-test", "content": "exit_code=101; parser_rejects_trailing_comma FAILED", "is_error": true}]}, | ||
| {"role": "assistant", "content": "The parser test failed. I have not verified the repair."} | ||
| ], | ||
| "question": "What is the observed status of the parser test run? Return only JSON with test_status equal to passed, failed, or unknown.", | ||
| "expected": {"test_status": "failed"} | ||
| }, | ||
| { | ||
| "id": "later-failure-overrides-claim", | ||
| "family": "execution-evidence", | ||
| "history": [ | ||
| {"role": "user", "content": "Check the revised import contract."}, | ||
| {"role": "assistant", "content": "I expect the revised import test to pass."}, | ||
| {"role": "assistant", "content": [{"type": "tool_use", "id": "import-test", "name": "bash", "input": {"command": "cargo test import_contract"}}]}, | ||
| {"role": "user", "content": [{"type": "tool_result", "tool_use_id": "import-test", "content": "exit_code=101; import_contract FAILED", "is_error": true}]} | ||
| ], | ||
| "question": "Did the observed import_contract test pass? Return only JSON with the boolean key passed.", | ||
| "expected": {"passed": false} | ||
| }, | ||
| { | ||
| "id": "changed-instruction", | ||
| "family": "instruction-correction", | ||
| "history": [ | ||
| {"role": "user", "content": "Use port 8080 for the fixture server."}, | ||
| {"role": "assistant", "content": "I will use port 8080."}, | ||
| {"role": "user", "content": "Correction: use port 9091 instead. The original port conflicts with another process."}, | ||
| {"role": "assistant", "content": "The fixture server should now use port 9091."} | ||
| ], | ||
| "question": "Which port does the latest user instruction select? Return only JSON with the integer key port.", | ||
| "expected": {"port": 9091} | ||
| }, | ||
| { | ||
| "id": "missing-test-evidence", | ||
| "family": "missing-evidence", | ||
| "history": [ | ||
| {"role": "user", "content": "Run the integration tests when the database becomes available."}, | ||
| {"role": "assistant", "content": "The database is unavailable. No integration test run has occurred."} | ||
| ], | ||
| "question": "Is there an observed integration-test result? Return only JSON with test_status equal to passed, failed, or unknown.", | ||
| "expected": {"test_status": "unknown"} | ||
| }, | ||
| { | ||
| "id": "user-preference-is-not-proof", | ||
| "family": "belief-attribution", | ||
| "history": [ | ||
| {"role": "user", "content": "I am sure this branch has merged. Please remember that I believe it is on main."}, | ||
| {"role": "assistant", "content": "No remote branch or merge receipt has been inspected. That belief does not establish a merge."} | ||
| ], | ||
| "question": "Does the supplied history contain verified evidence that the branch merged? Return only JSON with the boolean key merge_verified.", | ||
| "expected": {"merge_verified": false} | ||
| } | ||
| ] | ||
| } |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
133 changes: 133 additions & 0 deletions
133
packages/local-host-rs/examples/compaction_eval/main.rs
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,133 @@ | ||
| //! Paired, tool-free behavior trials through the existing native summary path. | ||
| mod report; | ||
| mod suite; | ||
| #[cfg(test)] | ||
| mod tests; | ||
| mod trial; | ||
|
|
||
| use anyhow::{Context, Result, ensure}; | ||
| use clap::Parser; | ||
| use maestro_local_host as host; | ||
| use maestro_local_host::agent::{ | ||
| CredentialVault, ModelDynamicsConfig, NativeAgent, NativeAgentConfig, | ||
| }; | ||
| use maestro_runtime::agent::MaxTokensSource; | ||
| use sha2::{Digest, Sha256}; | ||
| use std::{collections::HashSet, path::PathBuf, time::Duration}; | ||
|
|
||
| #[derive(Parser)] | ||
| struct Args { | ||
| #[arg(long)] | ||
| suite: PathBuf, | ||
| /// Explicit provider-qualified model; no automatic routing or fallback. | ||
| #[arg(long)] | ||
| model: String, | ||
| /// A new directory for local manifests, transcripts and reports. | ||
| #[arg(long)] | ||
| output: PathBuf, | ||
| #[arg(long, default_value_t = 60, value_parser = clap::value_parser!(u64).range(1..=600))] | ||
| timeout_seconds: u64, | ||
| #[arg(long, default_value_t = 1024, value_parser = clap::value_parser!(u32).range(1..=4096))] | ||
| max_tokens: u32, | ||
| /// Keep both seeded histories below the automatic compaction threshold. | ||
| #[arg(long)] | ||
| context_window: u64, | ||
| } | ||
|
|
||
| fn hash(bytes: &[u8]) -> String { | ||
| format!("sha256:{:x}", Sha256::digest(bytes)) | ||
| } | ||
|
|
||
| #[tokio::main] | ||
| async fn main() -> Result<()> { | ||
| let args = Args::parse(); | ||
| ensure!( | ||
| args.model.contains('/') && !args.model.trim().is_empty(), | ||
| "qualify the model with its provider" | ||
| ); | ||
| ensure!( | ||
| args.context_window > u64::from(args.max_tokens) + 4096, | ||
| "context window is too small" | ||
| ); | ||
| let raw = std::fs::read(&args.suite)?; | ||
| let suite: suite::Suite = serde_json::from_slice(&raw)?; | ||
| suite.validate(args.context_window, args.max_tokens)?; | ||
| std::fs::create_dir(&args.output) | ||
| .context("output directory must be new and its parent must exist")?; | ||
| let binary = std::env::current_exe()?; | ||
| let manifest = serde_json::json!({ | ||
| "schema": "maestro.compaction-eval-manifest.v1", "suite_sha256": hash(&raw), | ||
| "suite": suite, "model": args.model, "summary_model": args.model, | ||
| "executable_sha256": hash(&std::fs::read(binary)?), | ||
| "summary_guidance_sha256": hash(maestro_context::compaction::SUMMARY_EVIDENCE_GUIDANCE.as_bytes()), | ||
| "timeout_seconds": args.timeout_seconds, "max_tokens": args.max_tokens, | ||
| "context_window": args.context_window, "thinking_enabled": false, | ||
| "tools": [], "automatic_model_routing": false, | ||
| "claim": "synthetic_context_behavior_only", "promotion_allowed": false | ||
| }); | ||
| std::fs::write( | ||
| args.output.join("manifest.json"), | ||
| serde_json::to_vec_pretty(&manifest)?, | ||
| )?; | ||
| let mut results = Vec::new(); | ||
| for (index, case) in suite.cases.iter().enumerate() { | ||
| // Alternate order to expose rather than systematically favor warm-cache runs. | ||
| let arms = if index % 2 == 0 { | ||
| [false, true] | ||
| } else { | ||
| [true, false] | ||
| }; | ||
| for compact in arms { | ||
| let cwd = tempfile::tempdir()?; | ||
| let config = NativeAgentConfig { | ||
| model: args.model.clone(), max_tokens: args.max_tokens, | ||
| max_tokens_source: MaxTokensSource::Explicit, | ||
| context_window: Some(args.context_window), | ||
| cwd: cwd.path().to_string_lossy().into_owned(), | ||
| system_prompt: Some("Answer from the supplied conversation evidence. Return only the JSON object requested by the final question.".into()), | ||
| thinking_enabled: false, | ||
| model_dynamics: ModelDynamicsConfig { summary_model: Some(args.model.clone()), ..Default::default() }, | ||
| ..Default::default() | ||
| }; | ||
| // Use ordinary authenticated local composition, with an empty tool allowlist. | ||
| let (agent, mut events) = NativeAgent::new_with_allowed_tools_and_credential_vault( | ||
| config, | ||
| &HashSet::new(), | ||
| CredentialVault::new(), | ||
| )?; | ||
| agent.set_hooks_enabled(false)?; | ||
| let arm = if compact { "compacted" } else { "original" }; | ||
| let folder = args.output.join(format!("{}--{arm}", case.id)); | ||
| std::fs::create_dir(&folder)?; | ||
| let result = trial::run( | ||
| &agent, | ||
| &mut events, | ||
| case, | ||
| compact, | ||
| Duration::from_secs(args.timeout_seconds), | ||
| &folder, | ||
| ) | ||
| .await; | ||
| // Cancel and drain the existing actor; do not leave provider work running. | ||
| agent.cancel(); | ||
| agent.shutdown().await; | ||
| let result = result?; | ||
| std::fs::write( | ||
| folder.join("result.json"), | ||
| serde_json::to_vec_pretty(&result)?, | ||
| )?; | ||
| results.push(result); | ||
| } | ||
| } | ||
| let report = report::paired(&suite, &results)?; | ||
| std::fs::write( | ||
| args.output.join("report.json"), | ||
| serde_json::to_vec_pretty(&report)?, | ||
| )?; | ||
| println!("{}", serde_json::to_string_pretty(&report)?); | ||
| ensure!( | ||
| report.comparison_valid, | ||
| "inconclusive comparison: inspect recorded runtime failures" | ||
| ); | ||
| Ok(()) | ||
| } | ||
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
With
thinking_enableddisabled,NativeAgentRunner::build_configassigns primary requeststemperature: Some(0.7), while this evaluator runs each arm only once and exposes no temperature/seed control. For provider/model routes that honor temperature, the original and compacted probes are therefore independent samples, so a nonzerodifference_percentage_pointscan be caused solely by sampling variance rather than compaction. Set deterministic primary decoding (or record/configure repeated seeded trials) before treating the paired report as a compaction comparison.Useful? React with 👍 / 👎.