EXAMPLE DATA — fictional examples, not real model performance / 当前为示例数据
Today’s monitoring alerts (2026-10-08): 1 models triggered M3 and other rules. Details →

Rules

Rules version 2026-10-08.1 · Last updated 2026-10-08 · Rules changelog →

What this site is

In one sentence:A same-condition pelican archive for comparing versions and reasoning efforts — entertainment and reference only.

> A same-condition pelican archive for comparing versions and reasoning efforts — entertainment and reference only.

Tihu Pelican Board generates SVGs with one harness and one prompt so you can compare versions and efforts inside a family. It is not a ranking of overall model ability. In July 2026 Simon Willison wrote that the correlation “has been mostly severed now”. We take that as the brief: archive and family comparison.

The fixed prompt

In one sentence:Models always receive this English sentence. We never rewrite it.

> Models always receive this English sentence. We never rewrite it.

Generate an SVG of a pelican riding a bicycle

Chinese gloss (understanding only): 生成一张鹈鹕骑自行车的 SVG 图. Variant prompts get their own ids and boards.

How images are generated

In one sentence:All on the operator’s Tencent Cloud VPS, with the same cursor-agent, default settings, and a fresh empty folder.

> All on the operator’s Tencent Cloud VPS, with the same cursor-agent, default settings, and a fresh empty folder.

Each version × effort produces 5 images. Each run is an empty directory without `--force`. Each call times out after 900 seconds; failures retry up to 3 times at 30 / 120 / 600 seconds, and 5 consecutive sample failures pause the run. Tokens and cost are not available from cursor-agent. Results are not comparable to raw API runs. See `config/models.yaml`.

How long data is kept

In one sentence:Raw records stay on the server for 30 days. What the site shows is kept in the repo.

> Raw records stay on the server for 30 days. What the site shows is kept in the repo.

Daily raw records (full logs, model output, renders, gate and judge raw output, monitoring and alerts) are stored by date on the server and kept for 30 days, then deleted. Site-facing data (sample metadata, sanitized SVGs, scores, comments, monitoring and alert results) lives in this repo and stays published. Raw logs are not kept long-term; contact admin@tihu.app within 30 days if you need to check them.

Gate checks

In one sentence:Eight automatic checks run before judging. Failed images are shown but not scored.

> Eight automatic checks run before judging. Failed images are shown but not scored.

CodeMeaningThreshold
no_svgNo SVG / unparsable—
truncatedUnclosed or cut off—
embedded_raster`<image>` or data:image—
text_labelMostly texttext area > 0.2 or keywords pelican / bicycle / bike
external_refExternal href / url() / @import—
blankNearly blanknon-background < 0.005
off_canvasMostly outside the canvasvisible on canvas < 0.5
too_many_elementsToo large> 10000 elements or > 524288 bytes

Valid SVG rate = passes ÷ 5.

How AI judges score

In one sentence:They see the render only: blind description, yes/no checklist, then scores; two shuffled passes; same-vendor excluded.

> They see the render only: blind description, yes/no checklist, then scores; two shuffled passes; same-vendor excluded.

Three judges (Claude / GPT / Grok, all via cursor-agent). Batches of 3 images from different models; 2 order passes; parse failures retry 2 times. Same-vendor rows stay visible but do not count. Full prompts are in the disclosure below.

How scores are computed

In one sentence:The reference score is a weighted mean, shown to 1 decimal, sorted unrounded.

> The reference score is a weighted mean, shown to 1 decimal, sorted unrounded.

DimensionWeight
Pelican likeness0.25
Bicycle correctness0.25
Riding pose0.2
Composition0.15
Obvious errors (higher is better)0.15

A sample score averages counted judges and both passes. A variant score averages gate-passed batch samples. Fewer than 3 valid samples → insufficient. Best sample = highest AI reference score. The reference board defaults to each model’s default effort.

Voting and Elo rules

In one sentence:Phase 2 is coming soon; the rules are public now. The Elo column on the reference board currently says “coming soon”.

> Phase 2 is coming soon; the rules are public now. The Elo column on the reference board currently says “coming soon”.

RuleValue
Primary methodBradley-Terry, 0.95 CI
Elo (display only) initial rating1000
Elo K (games < 30 / ≥ 30)32 / 16
Minimum votes to appear on the visitor board15
Provisional below50
“Both bad”Counts as a tie 0.5

The visitor board and the AI reference score are independent: votes never change the AI score, and the AI score never changes votes.

Daily drift monitoring

In one sentence:One canary image per day versus baseline. Alerts add an “under monitoring” badge and never auto-change scores.

> One canary image per day versus baseline. Alerts add an “under monitoring” badge and never auto-change scores.

Schedule: 02:00 Asia/Shanghai. Cleared after 7 consecutive days without a trigger.

IdMeaningThresholdLevel
M1Consecutive failures2 dayswarning
M2Valid-rate drop30 pp below baseline, or 3 consecutive gate-fail dayswarning
M3Score dropbelow μ − max(1.5, 2σ), window 7 dayswarning
M4Duration outlierabove 2× or below 0.5× baselineinfo
M5Environment changerecorded immediatelyinfo
M6Model delistedrecorded immediatelyinfo

What this can and cannot show

In one sentence:It can show version/effort differences and whether the SVG is valid. It cannot measure overall ability across families.

> It can show version/effort differences and whether the SVG is valid. It cannot measure overall ability across families.

Can: same-harness version/effort differences; validity; rough duration. Cannot: overall ability or a ranking of model capability; comparison versus raw APIs or other sites; N=5 is small; judges are biased; cursor-agent has a hidden system prompt.

Disclaimer

Tihu Pelican Board is an independent hobby project, not affiliated with or endorsed by Simon Willison, his blog, pelicanbenchmark.com, or Pelican Zoo. The pelican-on-a-bicycle test was proposed by Simon Willison; we only cite his public posts with links.

Every image here was generated by this project with cursor-agent. We do not reuse third-party artwork. Results are affected by cursor-agent’s hidden system prompt and tools and cannot be compared directly with raw API runs.

The AI reference score is for entertainment and reference only. It is not an evaluation of any model or vendor. Names and trademarks belong to their owners.

Corrections, takedown, contact

In one sentence:Email a link and reason; we aim to reply in 7 days; outcomes go in the changelog.

> Email a link and reason; we aim to reply in 7 days; outcomes go in the changelog.

Contact: admin@tihu.app. Metadata and judge output: CC BY 4.0. SVGs follow vendor terms. Rules 2026-10-08.1, updated 2026-10-08.

Full judge prompts

Blind-description prompt

请用 2–3 句话客观描述这张图片里画了什么(有哪些物体、形状、颜色、相对位置)。
不要评价好坏,不要猜测它的用途或出处。
只输出 JSON:{"description": "……"}

Scoring prompt

你是一名严格、公正的插图评审。下面有若干张图片,分别标记为 A、B、C。
每张图片都是某个 AI 模型根据同一条提示词生成的 SVG 的渲染结果,提示词原文为:

"Generate an SVG of a pelican riding a bicycle"

你不知道每张图片来自哪个模型,请不要猜测,也不要因为图片顺序产生偏好。

请对每张图片按以下步骤评审:

第一步:客观描述。用 2–4 句话描述你在图中实际看到的内容(有哪些物体、形状、颜色、相对位置),不要评价。

第二步:逐项检查。对下列每一项只回答 true 或 false:
- has_beak_pouch:有长喙和喉囊
- two_wheels:有两个轮子
- frame:有连接两轮的车架
- on_seat:鸟坐在车座上
- feet_on_pedals:鸟脚在踏板上

第三步:按以下 5 个维度分别打 1–10 的整数分(10 为最好):
1. pelican(鹈鹕相似度):能否认出是鹈鹕——长喙、喉囊、身体与翅膀 / 腿的比例。
2. bicycle(自行车正确性):是否有两个轮子、车架、车把、踏板、链条,结构是否合理。
3. pose(骑行姿态):鹈鹕是否坐在车座上、脚是否在踏板上、翅膀是否握住车把。
4. composition(整体构图与美观):布局、配色、完整度、视觉吸引力。
5. errors(明显错误,分数越高表示错误越少):是否有部件漂浮、错误重叠、缺失、错位、画面被截断等问题;没有明显错误给 9–10 分。

第四步:用一句中文(不超过 40 个字)写出对该图最关键的点评。

评分要求:
- 只根据图片内容评分,不要考虑你认为的模型身份。
- 分数要拉开差距,不要所有维度都给中间分。
- 第三步的分数必须与第二步的检查结果一致(例如 two_wheels 为 false 时,bicycle 不应高于 4 分)。
- 如果图片空白或无法辨认,所有维度给 1 分,并在描述中说明。

请只输出如下 JSON,不要输出任何其他文字:
{
  "results": [
    {
      "label": "A",
      "description": "……",
      "checklist": { "has_beak_pouch": true, "two_wheels": true, "frame": true, "on_seat": true, "feet_on_pedals": true },
      "scores": { "pelican": 0, "bicycle": 0, "pose": 0, "composition": 0, "errors": 0 },
      "comment": "……"
    }
  ]
}

Judge panel

channel: cursor-agent · exclude_same_vendor: true

role vendor_id cursor_model_id Tie rate Order-inconsistency rate
claude anthropic <最新 Claude 型号> 0% 0%
gpt openai <最新 GPT 型号> 0% 0%
grok xai <最新 Grok 型号> 0% 0%

Same-vendor exclusion

Evaluated vendor Counted judges
anthropic gpt + grok
openai claude + grok
xai claude + gpt
any other vendor claude + gpt + grok