← All Test Cases
medium
CORE-002
v040 B core memory
Repetitions
1
Documents
0
Questions
0
Reasoning
DETERMINISTIC
v040
B2
core_memory
token_cap
📖 In Plain English
No description available yet for category "v040_B_core_memory".
⚙️ How a single rep runs
① Generate
Model creates 0 synthetic documents and 0 questions with unique canary tokens
→ Fresh content per run prevents memorization and proves real retrieval
② Ingest (MCP)
Model calls brain_ingest to store the 0 documents
→ Tests the brain's storage and indexing pipeline
③ Query (MCP)
Model answers the question using brain retrieval tools (search, fetch, context_pack, etc.)
→ Core test — does the brain return correct evidence and let the model build a faithful answer?
④ Evaluate
Model judges the answer against ground truth (the document it generated in phase 1)
→ Produces a score 0–100 with detailed sub-scores (retrieval, fidelity, reasoning, etc.)
This rep is run 1 times per test run. A pass requires score ≥ 85 and no critical failures.
🔬 Technical Instructions (raw prompts sent to AI)
Recent Run History
0 runsNo runs have included this test yet.
📄 Raw YAML cases/v040_B_core_memory/CORE-002.yaml
schema_version: "1.0" test_id: "CORE-002" category: "v040_B_core_memory" kind: "deterministic" severity: "medium" repetitions: 1 reasoning_type: "DETERMINISTIC" num_documents: 0 num_questions: 0 tags: ["v040", "B2", "core_memory", "token_cap"] assertion_module: "v040.B_core_memory.core_002" description: | Ingest 30 claims at importance=0.95; verify context_pack.core_memory is bounded by the ~500-token cap (typically 20-25 claims).