Back to All Work
LLM EvaluationLive Portal Active

Local Model Matrix

Enterprise-Grade Local LLM Discovery, Multi-Dimensional Filtering & Evaluation Suite

Local Model Matrix • Live Radar & Benchmark Matrix
LM Studio Ready
Local Model Matrix Interface Snapshot

System Simulator Sandbox

Node Ready
Sample Interaction:
Run Level 4 Olympiad Math & BFCL Function Calling benchmark against Qwen 30B vs Gemma 12B.
Execution Output:
Deterministic Verifier Output: • Level 4 AIME Math: 82.4% exact match (LaTeX symbolic verified) • BFCL Tool Calling: 94.1% schema adherence (Ajv validated) • TTFT: 142ms | TPS: 34.8 tok/s | Pareto Index: 9.2/10

The Problem

There are millions of open source models, but determining what model to use and for what is hard unless tested individually in applications. LMM is the first stage gate that filters and identifies local models that are worth your time for further validation and evaluation.

The Architecture Solution

Acts as a rapid first-stage filter and deterministic benchmark engine (sandboxed code execution, symbolic LaTeX solvers, Ajv Draft-07 schemas, IFEval rule engines) across a 5-Level Calibrated Stress Ladder with 7-Factor Failure Autopsies.

Key Innovations

  • First-Stage Gate Triage for thousands of open-source models
  • 11-Dimensional Capability Radar (Coding, Math, STEM, Instruction Following, Tool Calling, Vision, Med, Legal, Graphs, Speed)
  • Zero-Hallucination Deterministic Verification (No circular LLM judging LLM)
  • Hardware Efficiency Frontier & Pareto Optimization (Pass@1 Accuracy vs Tokens/Sec)
  • Real-Time Token Streaming with sub-millisecond TTFT and TPS profiling

Verified Telemetry

Capability Radar
11 dimensions
Evaluated Models
21 frontier models
Deterministic Verifiers
9 engines
Streaming TTFT
140ms
Assertion Pass Rate
99.9%

Tech Stack

Express.jsTypeScriptPrismaPostgreSQLLM Studio CLIAjvPython SubprocessSSE