◈ Competition Entry  ·  Visual Data Story

Beyond the Demo: 108 Assistants, 16 Real Tasks, One Honest Scorecard

108 assistants · 16 dimensions · one 1–10 scale · Overall is a running mean until fully tested. This infographic recreates the Scorecard matrix as granular RDF: every cell is a ScoreObservation linking one assistant to one dimension.

 ·  KG curated by kg-generator, rdf-infographic-skill, and Muse Spark 1.2 on behalf of Kingsley Idehen

Benchmark v0.2 · 15 tasks · 180 runs Updated 2026-09-15 · Last test 2026-09-14 Source assistantbenchmark.com Dimensions /dimensions Head-to-head /compare

Overview

Focus: comparisons and dimensions only — the benchmark's dimension definitions and the per-dimension Scorecard, typed for query and reuse.

What this collection does

Which assistants actually get real work done? This benchmark tests 108 AI assistants across 16 everyday dimensions — booking travel, handling email, remembering context, respecting permissions — and scores each 1–10 where the test applies. N/A means out-of-scope, — means not yet tested; Personality is judged from public quotes, never averaged.

Coverage: 0 Completed · 21 In progress · 32 Pending · 108 total · 5,565 public quotes. See How scoring works.

How to read the matrix

Desktop shows a semantic table with a sticky first column; phones show one card per assistant. Both carry identical facts. Every dimension and assistant name is a resolver link to its RDF IRI.

Comparison Matrix — Scorecard (granular)

TRANSPOSED: 21 assistants as rows · 16 dimensions as columns · Overall = running mean · Tested = scored / applicable. Colors encode score bands. Resolver links on every dimension and assistant.

AssistantOnline tasksBook a hotel stayTravelBook a flightPicksPick a restaurantPurchasingReorder on AmazonEmailReply to scheduling emailProactiveFlight day unpromptedRoutinesDaily digestConnected appsThree tools one requestPermissionsScoped access + hard ruleMemoryRecall preferencesPersonalityPublic opinion onlyPhone callsCall a businessGroup chatsPlan dinner in group chatMulti-stepFlight check-in chainRestraintKnow when not toImagesMake images/gamesOverallTestedSpeed
Muse9791010——109———————9.17/157s
Instinct81099989859———8——8.411/1520s
szn88788——771010—7—8—8.011/1539s
Pally789888—78—47979—7.613/15-
Ollie977888—78—49879—7.613/15-
Tomo79888——776——7—9—7.610/1411s
Asaply—N/A78N/A——————N/A——N/A—7.52/11-
Shuffle7.59988——7764—7—9—7.411/1455s
Grok Bot8—789——3————8—8—7.37/1510s
Caddy88777——7—103—7—8—7.210/1432s
Catch6—————7—————————6.52/1548s
Poke3—8—8777—8—582——6.310/1511s
Wajo / Fo277—6—7——8——————6.26/1533s
Asmi657585—68—43879—6.213/15-
Boba65858————64———7—6.18/15-
Folk666—————————————6.03/15-
Boski3—4—————6—————8—5.34/159s
tinyNature4—5—675——3——————5.06/159s
Brea3—7—————6—————4—5.04/157s
Town—————622————————3.33/157s
OpenInstinct——————1—————————1.01/1523s

Source table captured 2026-09-15 06:31 UTC from assistantbenchmark.com. Example high scores: Muse 10 on Purchasing/Email/Connected-apps, Instinct 10 on Travel/Memory, szn 10 on Memory/Personality.

Dimensions — the 16 tests

Each dimension is a Dimension (also schema:DefinedTerm) with its task prompt, ranking and tested count.

Online tasks

Carrying out an online task — Book a hotel stay. Completes a real browser or web workflow end to end.

20 tested · /dimensions/online_task

Travel

Travel booking — Book a flight and handle the trip, including check-in and changes.

14 tested

Purchasing

Reorder a product on Amazon — finds, compares and correctly stages a purchase.

12 tested

Email

Responding to emails — reply to a scheduling email in your voice.

14 tested

Proactive

Flight day, unprompted — acts or nudges usefully without being asked.

7 tested

Routines

Running a routine — daily digest for a week; reliable scheduled work.

7 tested

Connected apps

Third-party integrations — three tools, one request (calendar, inbox, Notion).

12 tested

Permissions

Permissions & privacy — scoped access, hard rule, clean revocation.

10 tested

Memory

Recall preferences and track context across the thread.

9 tested

Personality

Has a voice worth talking to — subjective, read from public quotes (not scored 1–10).

Public opinion only · 5,565 quotes

Phone calls

Call a business and get an answer — makes real phone calls and reports back.

7 tested

Group chats

Plan dinner in a group chat — works with several people at once.

4 tested

Multi-step

Chained tasks — strings several steps across tools into one job.

9 tested

Restraint

Proactive restraint — know when not to act; judgment on unprompted action.

5 tested

Images

Content creation / games — makes images, video or games on request.

11 tested

HowTo — Read & reuse the granular comparison

Every step below is a schema:HowToStep with an absolute IRI; headings are resolver-linked.

1

Open the source at assistantbenchmark.com and confirm Benchmark v0.2, 15 tasks, 180 runs, updated 2026-09-15.

Verify the bench-strip and Scorecard header — the provenance for all 336 cells.

2

Open /dimensions and read the 16 dimension cards to map each slug to its task.

Capture online_task, travel, recommendation_quality, purchasing, email_replies, proactive_behavior, running_routine, third_party_integrations, permissions_privacy, memory, personality, phone_calls, multiplayer_groups, chained_tasks, proactive_restraint, content_creation_games.

3

Pick the comparison scope: All 108 or the 21 in-progress detailed rows recreated here; note Overall and Tested.

Filter ScoreObservations by scoreStatus to separate scored vs N/A vs not_tested.

4

Read each cell as a ScoreObservation with forAssistant, forDimension, scoreValue, scoreStatus.

Query :ScoreObservation in Turtle/SPARQL to pivot by assistant or dimension.

5

Treat Personality as public-opinion-only; exclude it from means and compare via quotes.

Use opinion signals, not the 1–10 scale, for Personality.

6

Use /compare for pairwise head-to-head, or SPARQL delta over this graph.

Delta pattern: SELECT ?dim ?a ?b WHERE { ?o1 :forAssistant :assistant_muse ; :forDimension ?dim ; :scoreValue ?a . ?o2 :forAssistant :assistant_instinct ; :forDimension ?dim ; :scoreValue ?b }

FAQ

14 Questions — each is a schema:Question with a resolver link; answers are typed schema:Answer.

Assistant Benchmark at assistantbenchmark.com is a public scorecard for AI assistants you can text. Anchored to Scorecard, /dimensions, /compare. v0.2 · 108 assistants · 15 scored + Personality · 180 runs · updated 2026-09-15.
16 dimensions: 15 scored 1–10; Personality is public-opinion-only (5,565 quotes).
One published task per dimension — Book a hotel, Book a flight, Pick a restaurant, Reorder on Amazon, Reply to email, Flight-day nudge, Daily digest, Three tools, Scoped access, Recall preferences, Voice (quotes), Call a business, Group chat, Flight check-in chain, Know when not to act, Make images/games.
1–10 against anchors; — = not tested; N/A = doesn't apply. Overall is the running mean; Tested shows fraction like 7/15.
They are the In-progress set (21 assistants) from the Scorecard wide table as captured 2026-09-15 06:31 UTC, with exact per-dimension scores reproduced and typed as ScoreObservations.
Numeric = scored; N/A = not_applicable; — = not_tested. Distinct in Turtle for query filtering.
0 Completed, 21 In progress, 32 Pending (108 total). Displayed on the Scorecard bench-strip.
Use /compare or the SPARQL delta in the Workbench (FAQ 11).
From /dimensions cards: e.g. online_task 20, travel 14, recommendation_quality 20, etc.
List scored cells or compute Muse vs Instinct delta (see Workbench Q1/Q2).
Personality has no 1–10 audit; described as public opinion only in the dimension definition.
Source scorecard updated 2026-09-15 (last test 2026-09-14), v0.2, 180 runs. Each ScoreObservation is prov:wasDerivedFrom assistantbenchmark.com.
Every dimension and assistant IRI is reachable via linkeddata.uriburner.com/describe/?url={encodedIRI} for KG Explorer and table links.

Glossary

12 terms — each a schema:DefinedTerm in the DefinedTermSet.

Benchmark (Scorecard)

The Scorecard at assistantbenchmark.com — 108 assistants, v0.2, 180 runs, updated 2026-09-15.

Dimension

One of 16 tests; 15 scored 1–10, Personality is quotes-only. Each has one published task.

ScoreObservation

Reified cell linking assistant + dimension to scoreValue + scoreStatus.

Overall

Running mean until Tested reaches applicable-count denominator.

Tested

Fraction like 7/15 showing how many applicable dimensions are scored.

Knowledge Graph Explorer

Benchmark → Dimensions → Assistants → ScoreObservations (scored edges). Use Basic / Advanced, Classes / Predicates filters, and resolver-backed node/label links.

Graph

—
Modes
Drag to pan · scroll to zoom · drag node to pin · double-click to unpin · labels are resolver links

SPARQL Workbench

Live queries over the Turtle named graph via URIBurner (format=text/x-html+tr for SELECT; text/x-html-nice-turtle for CONSTRUCT). Each recipe projects an IRI alongside its labels.

Q1 — All scored cells (assistant, dimension, IRI + score)

PREFIX : <https://assistantbenchmark.com#>
PREFIX schema: <http://schema.org/>
SELECT ?assistantIri ?assistant ?dimensionIri ?dimension ?score WHERE {
  GRAPH <https://assistantbenchmark.com> {
    ?obs a :ScoreObservation ;
         :forAssistant ?assistantIri ;
         :forDimension ?dimensionIri ;
         :scoreStatus "scored" ;
         :scoreValue ?score .
    ?assistantIri schema:name ?assistant .
    ?dimensionIri schema:name ?dimension .
  }
} ORDER BY DESC(?score) LIMIT 50

Q2 — Muse vs Instinct delta

PREFIX : <https://assistantbenchmark.com#>
SELECT ?dimensionIri ?dimension ?muse ?instinct WHERE {
  GRAPH <https://assistantbenchmark.com> {
    ?o1 :forAssistant :assistant_muse ; :forDimension ?dimensionIri ; :scoreStatus "scored" ; :scoreValue ?muse .
    ?o2 :forAssistant :assistant_instinct ; :forDimension ?dimensionIri ; :scoreStatus "scored" ; :scoreValue ?instinct .
    ?dimensionIri schema:name ?dimension .
  }
}

Q3 — Dimensions by tested coverage

PREFIX : <https://assistantbenchmark.com#>
PREFIX schema: <http://schema.org/>
SELECT ?dimensionIri ?dimension ?tested (COUNT(?obs) AS ?scoredCells) WHERE {
  GRAPH <https://assistantbenchmark.com> {
    ?obs a :ScoreObservation ; :forDimension ?dimensionIri ; :scoreStatus "scored" .
    ?dimensionIri schema:name ?dimension ; :testedCount ?tested .
  }
} GROUP BY ?dimensionIri ?dimension ?tested ORDER BY DESC(?scoredCells)

Q4 — Top Overall with Tested denominator

PREFIX : <https://assistantbenchmark.com#>
PREFIX schema: <http://schema.org/>
SELECT ?assistantIri ?assistant ?overall ?tested WHERE {
  GRAPH <https://assistantbenchmark.com> {
    ?assistantIri a schema:SoftwareApplication ; schema:name ?assistant ; :overallScore ?overall ; :testedCount ?tested .
  }
} ORDER BY DESC(?overall) LIMIT 21

Endpoint: https://linkeddata.uriburner.com/sparql?default-graph-uri=&query={encoded}&format=text%2Fx-html%2Btr (SELECT) · format=text%2Fx-html-nice-turtle (CONSTRUCT/DESCRIBE). Named graph <https://assistantbenchmark.com>. IRIs are projected alongside labels so Virtuoso renders resolver links.