Overview
Focus: comparisons and dimensions only — the benchmark's dimension definitions and the per-dimension Scorecard, typed for query and reuse.
What this collection does
Which assistants actually get real work done? This benchmark tests 108 AI assistants across 16 everyday dimensions — booking travel, handling email, remembering context, respecting permissions — and scores each 1–10 where the test applies. N/A means out-of-scope, — means not yet tested; Personality is judged from public quotes, never averaged.
Coverage: 0 Completed · 21 In progress · 32 Pending · 108 total · 5,565 public quotes. See How scoring works.
How to read the matrix
Desktop shows a semantic table with a sticky first column; phones show one card per assistant. Both carry identical facts. Every dimension and assistant name is a resolver link to its RDF IRI.
Comparison Matrix — Scorecard (granular)
TRANSPOSED: 21 assistants as rows · 16 dimensions as columns · Overall = running mean · Tested = scored / applicable. Colors encode score bands. Resolver links on every dimension and assistant.
| Assistant | Online tasks | Travel | Picks | Purchasing | Proactive | Routines | Connected apps | Permissions | Memory | Personality | Phone calls | Group chats | Multi-step | Restraint | Images | Overall | Tested | Speed | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Muse | 9 | 7 | 9 | 10 | 10 | — | — | 10 | 9 | — | — | — | — | — | — | — | 9.1 | 7/15 | 7s |
| Instinct | 8 | 10 | 9 | 9 | 9 | 8 | 9 | 8 | 5 | 9 | — | — | — | 8 | — | — | 8.4 | 11/15 | 20s |
| szn | 8 | 8 | 7 | 8 | 8 | — | — | 7 | 7 | 10 | 10 | — | 7 | — | 8 | — | 8.0 | 11/15 | 39s |
| Pally | 7 | 8 | 9 | 8 | 8 | 8 | — | 7 | 8 | — | 4 | 7 | 9 | 7 | 9 | — | 7.6 | 13/15 | - |
| Ollie | 9 | 7 | 7 | 8 | 8 | 8 | — | 7 | 8 | — | 4 | 9 | 8 | 7 | 9 | — | 7.6 | 13/15 | - |
| Tomo | 7 | 9 | 8 | 8 | 8 | — | — | 7 | 7 | 6 | — | — | 7 | — | 9 | — | 7.6 | 10/14 | 11s |
| Asaply | — | N/A | 7 | 8 | N/A | — | — | — | — | — | — | N/A | — | — | N/A | — | 7.5 | 2/11 | - |
| Shuffle | 7.5 | 9 | 9 | 8 | 8 | — | — | 7 | 7 | 6 | 4 | — | 7 | — | 9 | — | 7.4 | 11/14 | 55s |
| Grok Bot | 8 | — | 7 | 8 | 9 | — | — | 3 | — | — | — | — | 8 | — | 8 | — | 7.3 | 7/15 | 10s |
| Caddy | 8 | 8 | 7 | 7 | 7 | — | — | 7 | — | 10 | 3 | — | 7 | — | 8 | — | 7.2 | 10/14 | 32s |
| Catch | 6 | — | — | — | — | — | 7 | — | — | — | — | — | — | — | — | — | 6.5 | 2/15 | 48s |
| Poke | 3 | — | 8 | — | 8 | 7 | 7 | 7 | — | 8 | — | 5 | 8 | 2 | — | — | 6.3 | 10/15 | 11s |
| Wajo / Fo | 2 | 7 | 7 | — | 6 | — | 7 | — | — | 8 | — | — | — | — | — | — | 6.2 | 6/15 | 33s |
| Asmi | 6 | 5 | 7 | 5 | 8 | 5 | — | 6 | 8 | — | 4 | 3 | 8 | 7 | 9 | — | 6.2 | 13/15 | - |
| Boba | 6 | 5 | 8 | 5 | 8 | — | — | — | — | 6 | 4 | — | — | — | 7 | — | 6.1 | 8/15 | - |
| Folk | 6 | 6 | 6 | — | — | — | — | — | — | — | — | — | — | — | — | — | 6.0 | 3/15 | - |
| Boski | 3 | — | 4 | — | — | — | — | — | 6 | — | — | — | — | — | 8 | — | 5.3 | 4/15 | 9s |
| tinyNature | 4 | — | 5 | — | 6 | 7 | 5 | — | — | 3 | — | — | — | — | — | — | 5.0 | 6/15 | 9s |
| Brea | 3 | — | 7 | — | — | — | — | — | 6 | — | — | — | — | — | 4 | — | 5.0 | 4/15 | 7s |
| Town | — | — | — | — | — | 6 | 2 | 2 | — | — | — | — | — | — | — | — | 3.3 | 3/15 | 7s |
| OpenInstinct | — | — | — | — | — | — | 1 | — | — | — | — | — | — | — | — | — | 1.0 | 1/15 | 23s |
Source table captured 2026-09-15 06:31 UTC from assistantbenchmark.com. Example high scores: Muse 10 on Purchasing/Email/Connected-apps, Instinct 10 on Travel/Memory, szn 10 on Memory/Personality.
Dimensions — the 16 tests
Each dimension is a Dimension (also schema:DefinedTerm) with its task prompt, ranking and tested count.
Online tasks
Carrying out an online task — Book a hotel stay. Completes a real browser or web workflow end to end.
20 tested · /dimensions/online_taskTravel
Travel booking — Book a flight and handle the trip, including check-in and changes.
14 testedPicks (Recommendation quality)
Pick a restaurant with constraints — relevance, taste and constraint-following.
20 testedPurchasing
Reorder a product on Amazon — finds, compares and correctly stages a purchase.
12 testedResponding to emails — reply to a scheduling email in your voice.
14 testedProactive
Flight day, unprompted — acts or nudges usefully without being asked.
7 testedRoutines
Running a routine — daily digest for a week; reliable scheduled work.
7 testedConnected apps
Third-party integrations — three tools, one request (calendar, inbox, Notion).
12 testedPermissions
Permissions & privacy — scoped access, hard rule, clean revocation.
10 testedMemory
Recall preferences and track context across the thread.
9 testedPersonality
Has a voice worth talking to — subjective, read from public quotes (not scored 1–10).
Public opinion only · 5,565 quotesPhone calls
Call a business and get an answer — makes real phone calls and reports back.
7 testedGroup chats
Plan dinner in a group chat — works with several people at once.
4 testedMulti-step
Chained tasks — strings several steps across tools into one job.
9 testedRestraint
Proactive restraint — know when not to act; judgment on unprompted action.
5 testedImages
Content creation / games — makes images, video or games on request.
11 testedHowTo — Read & reuse the granular comparison
Every step below is a schema:HowToStep with an absolute IRI; headings are resolver-linked.
Open the source at assistantbenchmark.com and confirm Benchmark v0.2, 15 tasks, 180 runs, updated 2026-09-15.
Verify the bench-strip and Scorecard header — the provenance for all 336 cells.
Open /dimensions and read the 16 dimension cards to map each slug to its task.
Capture online_task, travel, recommendation_quality, purchasing, email_replies, proactive_behavior, running_routine, third_party_integrations, permissions_privacy, memory, personality, phone_calls, multiplayer_groups, chained_tasks, proactive_restraint, content_creation_games.
Pick the comparison scope: All 108 or the 21 in-progress detailed rows recreated here; note Overall and Tested.
Filter ScoreObservations by scoreStatus to separate scored vs N/A vs not_tested.
Read each cell as a ScoreObservation with forAssistant, forDimension, scoreValue, scoreStatus.
Query :ScoreObservation in Turtle/SPARQL to pivot by assistant or dimension.
Treat Personality as public-opinion-only; exclude it from means and compare via quotes.
Use opinion signals, not the 1–10 scale, for Personality.
Use /compare for pairwise head-to-head, or SPARQL delta over this graph.
Delta pattern: SELECT ?dim ?a ?b WHERE { ?o1 :forAssistant :assistant_muse ; :forDimension ?dim ; :scoreValue ?a . ?o2 :forAssistant :assistant_instinct ; :forDimension ?dim ; :scoreValue ?b }
Cite with provenance: assistantbenchmark.com + dimensions/slug + this Turtle's named graph.
Include prov:wasDerivedFrom and dct:source links.
FAQ
14 Questions — each is a schema:Question with a resolver link; answers are typed schema:Answer.
scored; N/A = not_applicable; — = not_tested. Distinct in Turtle for query filtering.linkeddata.uriburner.com/describe/?url={encodedIRI} for KG Explorer and table links.Glossary
12 terms — each a schema:DefinedTerm in the DefinedTermSet.
Benchmark (Scorecard)
The Scorecard at assistantbenchmark.com — 108 assistants, v0.2, 180 runs, updated 2026-09-15.
Dimension
One of 16 tests; 15 scored 1–10, Personality is quotes-only. Each has one published task.
ScoreObservation
Reified cell linking assistant + dimension to scoreValue + scoreStatus.
Overall
Running mean until Tested reaches applicable-count denominator.
Tested
Fraction like 7/15 showing how many applicable dimensions are scored.
N/A — Does not apply
Dimension out of scope for that product.
— — Not tested yet
Dimension applies but no run logged yet.
Proactive restraint
Dimension 15: judgment on when not to act.
Permissions & privacy
Dimension 9: scoped access and clean revocation.
Third-party integrations
Dimension 8: three tools, one request.
Head-to-head
Paired view at /compare.
Public quotes
5,565 quotes backing Personality.
Knowledge Graph Explorer
Benchmark → Dimensions → Assistants → ScoreObservations (scored edges). Use Basic / Advanced, Classes / Predicates filters, and resolver-backed node/label links.
Graph
—SPARQL Workbench
Live queries over the Turtle named graph via URIBurner (format=text/x-html+tr for SELECT; text/x-html-nice-turtle for CONSTRUCT). Each recipe projects an IRI alongside its labels.
Q1 — All scored cells (assistant, dimension, IRI + score)
PREFIX : <https://assistantbenchmark.com#>
PREFIX schema: <http://schema.org/>
SELECT ?assistantIri ?assistant ?dimensionIri ?dimension ?score WHERE {
GRAPH <https://assistantbenchmark.com> {
?obs a :ScoreObservation ;
:forAssistant ?assistantIri ;
:forDimension ?dimensionIri ;
:scoreStatus "scored" ;
:scoreValue ?score .
?assistantIri schema:name ?assistant .
?dimensionIri schema:name ?dimension .
}
} ORDER BY DESC(?score) LIMIT 50Q2 — Muse vs Instinct delta
PREFIX : <https://assistantbenchmark.com#>
SELECT ?dimensionIri ?dimension ?muse ?instinct WHERE {
GRAPH <https://assistantbenchmark.com> {
?o1 :forAssistant :assistant_muse ; :forDimension ?dimensionIri ; :scoreStatus "scored" ; :scoreValue ?muse .
?o2 :forAssistant :assistant_instinct ; :forDimension ?dimensionIri ; :scoreStatus "scored" ; :scoreValue ?instinct .
?dimensionIri schema:name ?dimension .
}
}Q3 — Dimensions by tested coverage
PREFIX : <https://assistantbenchmark.com#>
PREFIX schema: <http://schema.org/>
SELECT ?dimensionIri ?dimension ?tested (COUNT(?obs) AS ?scoredCells) WHERE {
GRAPH <https://assistantbenchmark.com> {
?obs a :ScoreObservation ; :forDimension ?dimensionIri ; :scoreStatus "scored" .
?dimensionIri schema:name ?dimension ; :testedCount ?tested .
}
} GROUP BY ?dimensionIri ?dimension ?tested ORDER BY DESC(?scoredCells)Q4 — Top Overall with Tested denominator
PREFIX : <https://assistantbenchmark.com#>
PREFIX schema: <http://schema.org/>
SELECT ?assistantIri ?assistant ?overall ?tested WHERE {
GRAPH <https://assistantbenchmark.com> {
?assistantIri a schema:SoftwareApplication ; schema:name ?assistant ; :overallScore ?overall ; :testedCount ?tested .
}
} ORDER BY DESC(?overall) LIMIT 21Endpoint: https://linkeddata.uriburner.com/sparql?default-graph-uri=&query={encoded}&format=text%2Fx-html%2Btr (SELECT) · format=text%2Fx-html-nice-turtle (CONSTRUCT/DESCRIBE). Named graph <https://assistantbenchmark.com>. IRIs are projected alongside labels so Virtuoso renders resolver links.