Leaderboard & Hardware Efficiency Diagnostic
Deterministic evaluation of Small Language Models (1B–3.8B) vs Workstation (7B–9B) as read-only DevOps incident copilots.
Safety audit under rootless isolated sandbox (--read-only, --net=none, --tmpfs).
Hardware Provenance:
1x NVIDIA RTX 4090 24GB • Ollama CUDA 12.8 • RunPod Reference Node
100% Standalone (No-CORS)
Top Benchmark
—
—
RAM Sweet Spot
1.1 – 2.5 GB
Micro-Edge Floor ≤ 4 GB
Top Median TTFT
—
Reactive token-0 streaming
Zero-Catastrophe Rate
100.0%
Zero destructive commands
Audited Models
12
—
The Pareto Frontier Max Efficiency
SafeOps Index (Y) vs Peak RAM in GB (X). The green curve marks non-dominated models.
• Bastion Threshold: 4.0 GB RAM
Memory Penalty: (4.0 / RAM_peak)^0.5
Responsiveness vs Factual Precision TTFT (ms)
Diagnostic precision (%) as a function of streaming first-token latency.
• Tooltip: hover for throughput (tok/s)
Ideal target: top-left corner (<100 ms, 100%)
Competency Profile by DevOps Axis 4 Pillars
RCA (Diagnostic) • Blast Radius (Safety) • Surgical Diff (Non-regression) • Sanity Check (Dry-run).
Display:
Official Leaderboard
(12 models)
| # | Model | Division | SafeOps Index | Precision | Safety | Halluc./1k | Peak RAM | TTFT | Throughput | Detail |
|---|
Click any row to inspect multi-axis breakdown, flags audit, and traces.
Interactive sorting on all columns • 🔗 to copy recruiter link