Methodology
A transparent, independent framework for scoring AI platforms from 0 to 100. Seven weighted categories, public data sources, and auditable math — no black boxes, no pay-to-play.
Composite Formula
Steiner Index = Σ (category score × category weight)
Each category scored 0–100 · weights sum to 100% · composite scaled 0–100
Principles
Every weight, data source, and scoring rule is public. No black boxes, no proprietary algorithms — just open, auditable math.
No platform can pay to influence its score. Commercial relationships are publicly disclosed on each platform’s score page.
A visible methodology changelog tracks every change to weights or criteria, so historical scores remain comparable over time.
Scored platforms get a public right of reply. Rebuttals appear alongside the score, on the same page.
Calculation
Every platform receives a 0–100 score in each of the seven categories, drawn from public data sources and independent verification.
Each category score is multiplied by its published weight (expressed as a decimal). Weights are fixed, public, and sum to 100%.
The weighted scores are summed into a single composite — the Steiner Index — on the same 0–100 scale.
The composite ships with a visible last-updated date and source citations per cell, plus a methodology changelog for auditability.
Worked example
If a platform scores 92 in Capability (25%), 85 in Reliability (15%), … the composite is:
(92×0.25) + (85×0.15) + (75×0.15) + (65×0.10) + (95×0.10) + (95×0.10) + (95×0.15) = 87
Categories
Output quality, task accuracy, benchmark performance relative to category peers
Data sources
Published benchmark results, independent evaluations, structured internal testing.
Scoring approach
Output quality, task accuracy, and benchmark performance relative to category peers. Requires independent verification before going live.
Historical uptime, latency consistency, incident frequency and transparency
Data sources
Public status pages, third-party monitoring services.
Scoring approach
Historical uptime, latency consistency, incident frequency, and outage transparency. Quiet outages score lower than publicly reported ones.
Data handling policy, training-data opt-outs, security certifications, transparency reporting
Data sources
Published data-handling policies, security certifications (SOC 2, ISO 27001), transparency reporting.
Scoring approach
Data handling, training-data opt-outs, certifications, and disclosure. Weight likely to rise as AI regulation develops.
Clarity of pricing, hidden costs, value relative to capability tier
Data sources
Public pricing pages, tier documentation, terms of service.
Scoring approach
Clarity of pricing, hidden costs, and value relative to capability tier. Confusing multi-tier structures are docked.
API quality, third-party integrations, developer tooling
Data sources
API documentation, integration directories, developer tooling audits.
Scoring approach
API quality, third-party integrations, and developer tooling depth.
Frequency and substance of meaningful updates, responsiveness to user feedback
Data sources
Changelogs, release notes, product announcements.
Scoring approach
Frequency and substance of meaningful updates, plus responsiveness to user feedback.
Documentation quality, support responsiveness, community size and health
Data sources
Support ticket benchmarks, forum & Discord activity, documentation audits.
Scoring approach
Documentation quality, support responsiveness, and community size and health.
Cadence
Capability & Performance
Monthly
Fast-moving; new models and benchmarks land constantly.
Reliability & Uptime
Monthly
Status data updates continuously.
Data Privacy & Trust
Quarterly
Policies and certifications change slowly.
Pricing Transparency
Monthly
Tier structures and prices shift often.
Ecosystem & Integration
Quarterly
API and integration surface evolves gradually.
Update Velocity
Monthly
Release cadence is measurable in real time.
Community & Support
Quarterly
Community health shifts on longer cycles.
Governance
No platform can pay to influence its score. Any commercial relationship between Steiner Index and a scored platform must be publicly disclosed on that platform’s score page.
Scores are refreshed on a rolling basis — monthly for fast-moving categories like capability, quarterly for slower-moving categories like privacy policy. A visible methodology changelog tracks any change to weights or criteria, so historical scores remain comparable and auditable.
Disputed scores get a public right of reply. Platforms can submit a rebuttal that appears alongside their score. If they think we got it wrong, they can say so — right there on the same page.
A benchmark that’s afraid of being challenged isn’t a benchmark. It’s a marketing tool. We’re building something different.