A research-grade backtesting system for comparing trading policies (baselines, LLM, LangGraph debate) with strict scientific rigor.
- Deterministic simulation: Same config + seed = same results
- No look-ahead: News
available_at <= sim_timestrictly enforced - Pluggable policies: Protocol-based design for easy extensibility
- Decision logging: Full audit trail for post-hoc analysis
- Config-driven experiments: Pydantic schemas for reproducibility
trading_lab/ # Simulation & policies
|-- core/ # Types, scheduler, broker, engine
|-- policies/ # Trading policies
| |-- baselines/ # Simple strategies (buy-hold, momentum)
| |-- llm/ # Single LLM policy (M+2)
| `-- agentic/ # LangGraph-based agentic policy (M+3)
|-- data/ # Data loading for sim
|-- experiments/ # Config & runner
|-- storage/ # Postgres & Parquet
|-- evaluation/ # Metrics & analysis
`-- tests/ # Contract tests
data_preparation/ # Infrastructure (separate package)
data/ # Shared data (Parquet files)
# Install dependencies
pip install -r requirements.txt
# Set up environment variables (optional but recommended)
cp env.example .env
# Edit .env with your API keys
# Run contract tests
pytest trading_lab/tests/ -vSome features require additional packages:
# For database storage (PostgreSQL)
pip install psycopg2-binary
# For agentic policy (LangGraph)
pip install langgraph>=0.2.0The system uses environment variables for API keys and database configuration. Create a .env file from the template:
cp env.example .envThen edit .env with your actual values:
- OPENAI_API_KEY: For GPT-5.2, GPT-5 Nano, GPT-5 Mini, GPT-4o, o1, o3-mini
- OPENROUTER_API_KEY: For accessing multiple LLM providers via OpenRouter
- DB_* variables**: PostgreSQL connection (defaults match docker-compose.yml)
The .env file is automatically loaded when using the CLI (python -m trading_lab). For library usage, you can either:
- Set environment variables manually
- Use
python-dotenv:from dotenv import load_dotenv; load_dotenv()
| Type | Purpose |
|---|---|
Bar |
OHLCV bar (ts = end time, UTC) |
NewsEvent |
News with published_at and available_at |
ActionBundle |
Policy output (list of OrderIntent + rationale) |
Observation |
What policy sees (bars, news, portfolio) |
DecisionRecord |
Audit trail with forward returns |
Bar.ts= bar END time (UTC)- Market orders fill at next bar's OPEN
- All timestamps must be timezone-aware (UTC)
The system supports configurable execution costs including fees and slippage, with optional Amihud Illiquidity Measure (2002) adjustment for scientifically-grounded liquidity-based slippage.
- Two-part fees: Fixed fee + percentage fee
- Fixed slippage: Simple basis points model (default)
- Amihud-adjusted slippage: Volatility and volume-based adjustment (optional)
- Post-processing: Recalculate costs with different parameters without re-running simulations
from trading_lab.experiments.config import ExecutionConfig
# Fixed costs (default)
config = ExecutionConfig(
fee_per_trade=1.0,
fee_pct=0.001,
base_slippage_bps=5.0
)
# With Amihud adjustment
config = ExecutionConfig(
base_slippage_bps=5.0,
use_amihud_adjustment=True,
amihud_lookback=20,
market_amihud_baseline=1e-8,
amihud_sensitivity=0.5
)See EXECUTION_COSTS.md for detailed documentation, formulas, and academic references.
The agentic policy uses a LangGraph-based architecture with specialized nodes:
| Node | Role | Mode |
|---|---|---|
| Trading Agent (PM) | Scans market/news, selects candidate tickers | LLM-powered |
| Analyst | Evaluates candidates with direction/confidence | LLM-powered (single mode) |
| Risk Manager | Constructs portfolio with sizing and constraints | Deterministic |
from trading_lab.policies.agentic import AgenticPolicy, AgenticGraphConfig
config = AgenticGraphConfig(
trading_agent=TradingAgentConfig(
top_k_candidates=5,
objective="news_driven",
),
analyst=AnalystConfig(mode="single"),
risk_manager=RiskManagerConfig(
sizing_method="confidence_weighted",
max_position_pct=0.20,
),
valid_tickers=["AAPL", "GOOGL", "MSFT"],
)
policy = AgenticPolicy(config)
action = policy.act(observation)For the agentic policy, install LangGraph:
pip install langgraph>=0.2.0Note: The agentic policy includes a fallback executor that works without LangGraph installed.
When you run experiments via the ExperimentRunner, each repetition creates a run directory under results/:
results/<experiment_name>/<run_id>/config.json– full experiment configuration for that runequity.csv– equity curve over timedecisions.jsonl– high-level decision logfills.jsonl– execution log (if enabled)metrics.json– summary performance metricsllm_calls/– one JSON file per LLM request/response
The llm_calls/ folder contains files like:
run_0001__node_trading_agent__attempt_1.jsonrun_0002__node_analyst_team__attempt_1__tool_0.json
Each file is a single JSON object with:
- Metadata:
run_id,experiment_name,node_name,call_index,attempt, optionaltool_iteration,timestamp - Request: provider, model, output mode, prompt version, temperature,
max_tokens, schema name (if any), full chatmessages, and tools (for tool-using calls) - Response: raw model
contentplus token usage, latency, finish reason, and error (if any)
This structure makes it easy to inspect and analyze every LLM interaction alongside decisions and performance for full reproducibility.
The project includes an interactive Streamlit dashboard for exploring and analyzing experiment results.
streamlit run trading_lab/frontend/app.py| Page | Description |
|---|---|
| Experiment Analysis | Browse and filter experiments from the database |
| Strategy Overview | High-level comparison of all strategies |
| Performance Analysis | Equity curves, drawdowns, and risk-return metrics |
| Statistical Tests | Hypothesis testing and significance analysis |
| Decision Quality | Forward returns and decision timing analysis |
| Trading Patterns | Entry/exit patterns and holding periods |
| Cost Analysis | Fee/slippage breakdown and cost efficiency |
| Agentic Insights | Debate effectiveness and memory evolution (for agentic policies) |
| Conclusions | Interactive conclusion builder for generating reports |
- Date filtering: Focus analysis on specific time periods
- Interactive charts: Zoom, pan, and hover for details (Plotly)
- Export capabilities: Generate markdown/HTML reports
- Database integration: Load results directly from PostgreSQL
The system is highly configurable through YAML experiment files. Key customizable areas:
| Area | What You Can Change |
|---|---|
| Dataset | Tickers, date range, interval (1m/5m/15m/1h/1d), news inclusion |
| Simulation | Initial capital, scheduler type, lookback bars, max position size |
| Execution Costs | Fees (fixed + %), slippage, Amihud adjustment parameters |
| Policy | Strategy type, LLM model/provider, temperature, caching |
| Agentic Graph | Node configs (Trading Agent, Analyst, Risk Manager), debate settings |
| Storage | Database vs file output, connection settings |