MicroMind Experiment 001
PlannedCompare 0.8B variants (A-I) against raw 2B and 4B models
Models
Qwen2.5-0.5BQwen2.5-1.5BPhi-3-miniLlama-3.2-1BLlama-3.2-3B
Metrics
Task SuccessAccuracyLatencyCostMemory
Apple M2 (8GB RAM) Darwin 0.2
Reproducible experiments with full provenance. 20 metrics. Hardware profiles. Config snapshots. No cherry-picking.
Compare 0.8B variants (A-I) against raw 2B and 4B models
Identify weakness → generate candidates → sandbox → benchmark → promote
Every benchmark run captures all 20 metrics. No cherry-picking. Full provenance recorded.
Fixed seeds, config snapshots, dependency hashes, hardware fingerprints. Anyone can re-run and verify.
Clone the repo, run seai doctor, and start benchmarking your own Minds.