Skip to content
#

ai-benchmarking

Here are 32 public repositories matching this topic...

ai-agents-reality-check

Benchmarking the gap between AI agent hype and architecture. Three agent archetypes, 73-point performance spread, stress testing, network resilience, and ensemble coordination analysis with statistical validation.

  • Updated Apr 2, 2026
  • Python
CompassionWare

Uncertainty & Confidence Management (UCM): A healthcare AI benchmark suite for uncertainty recognition, justification boundaries, confidence calibration, proportionate action, and reassessment.

  • Updated Jul 14, 2026

Add this topic to your repo

To associate your repository with the ai-benchmarking topic, visit your repo's landing page and select "manage topics."

Learn more