Leaderboard & Hardware Efficiency Diagnostic

Deterministic evaluation of Small Language Models (1B–3.8B) vs Workstation (7B–9B) as read-only DevOps incident copilots. Safety audit under rootless isolated sandbox (--read-only, --net=none, --tmpfs).

Hardware Provenance: 1x NVIDIA RTX 4090 24GB • Ollama CUDA 12.8 • RunPod Reference Node
100% Standalone (No-CORS)
Top Benchmark
RAM Sweet Spot
1.1 – 2.5 GB
Micro-Edge Floor ≤ 4 GB
Top Median TTFT
Reactive token-0 streaming
Zero-Catastrophe Rate
100.0%
Zero destructive commands
Audited Models
12

The Pareto Frontier Max Efficiency

SafeOps Index (Y) vs Peak RAM in GB (X). The green curve marks non-dominated models.

• Bastion Threshold: 4.0 GB RAM Memory Penalty: (4.0 / RAM_peak)^0.5

Responsiveness vs Factual Precision TTFT (ms)

Diagnostic precision (%) as a function of streaming first-token latency.

• Tooltip: hover for throughput (tok/s) Ideal target: top-left corner (<100 ms, 100%)

Competency Profile by DevOps Axis 4 Pillars

RCA (Diagnostic) • Blast Radius (Safety) • Surgical Diff (Non-regression) • Sanity Check (Dry-run).

Display:
Official Leaderboard (12 models)
# Model Division SafeOps Index Precision Safety Halluc./1k Peak RAM TTFT Throughput Detail
Click any row to inspect multi-axis breakdown, flags audit, and traces. Interactive sorting on all columns • 🔗 to copy recruiter link