P PPTArena
API keys

Keys stay in your browser's local storage and are only sent with requests you run.

The first benchmark for editing real PowerPoint decks.

Michael Ofengenden · Yunze Man · Ziqi Pang · Liang-Yan Gui · Yu-Xiong Wang

100 Real decks
2,125 Slides
1,300+ Paired edits
IF + VQ Judge metrics

Leaderboard

Matched 25-case hard subset used for cost-sensitive agent comparisons. 10 systems · 25 cases · higher is better

PPTArena score

Combined instruction following + visual quality, normalized to 100

57.7
53.7
53.4
53.2
52.6
44.7
43.4
43.2
33.2
29.6
Claude (CUA)
Codex (GPT-5.5 xhigh)
Claude Code (Opus 4.8)
Gemini CLI (3.5 Flash)
G OpenCode (GLM-5.2)
OpenCode (Kimi K2.7 Code)
ChatGPT Agent (CUA)
D OpenCode (DeepSeek V4 Pro)
OpenCode (MiniMax-M3)
MiniMax Agent (CUA)
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Claude (CUA) Anthropic · Claude 3.7 Sonnet
57.7 65.6 49.9 25/25 57.7 Kimi K2.6
2
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
53.7 53.2 54.2 25/25 53.7 4.8 min Kimi K2.6
3
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
53.4 56.7 50.1 25/25 53.4 2.3 min Kimi K2.6
4
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
53.2 59.1 47.2 25/25 53.2 4.2 min Kimi K2.6
5
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
52.6 57.6 47.6 25/25 52.6 12.0 min Kimi K2.6
6
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
44.7 44.3 45.1 25/25 44.7 12.2 min Kimi K2.6
7
ChatGPT Agent (CUA) OpenAI · ChatGPT Agent
43.4 46.1 40.6 25/25 43.4 Kimi K2.6
8
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
43.2 44.4 42.1 25/25 43.2 6.7 min Kimi K2.6
9
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
33.2 35.1 31.4 25/25 33.2 10.1 min Kimi K2.6
10
MiniMax Agent (CUA) MiniMax · MiniMax Agent
29.6 30.5 28.6 25/25 29.6 Kimi K2.6
IF = instruction following, VQ = visual quality; each judged 0–5 per case by a VLM judge. Scores map to a percentage on a smooth curve where 5/5 = 100% and 4/5 ≈ 92% (a perfect score is the true 100%, with diminishing returns near the top); missing/unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.
Cases tagged for content and semantic editing. 10 systems · 20 cases · higher is better

PPTArena score

Combined instruction following + visual quality, normalized to 100

56.5
55.0
52.0
50.2
49.9
46.2
42.3
42.2
33.9
28.0
Claude (CUA)
Codex (GPT-5.5 xhigh)
G OpenCode (GLM-5.2)
Claude Code (Opus 4.8)
Gemini CLI (3.5 Flash)
OpenCode (Kimi K2.7 Code)
D OpenCode (DeepSeek V4 Pro)
ChatGPT Agent (CUA)
OpenCode (MiniMax-M3)
MiniMax Agent (CUA)
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Claude (CUA) Anthropic · Claude 3.7 Sonnet
56.5 65.6 47.5 20/20 56.5 Kimi K2.6
2
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
55.0 55.8 54.3 20/20 55.0 4.8 min Kimi K2.6
3
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
52.0 58.7 45.3 20/20 52.0 12.0 min Kimi K2.6
4
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
50.2 56.1 44.2 20/20 50.2 2.3 min Kimi K2.6
5
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
49.9 56.9 42.8 20/20 49.9 4.2 min Kimi K2.6
6
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
46.2 46.2 46.3 20/20 46.2 12.2 min Kimi K2.6
7
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
42.3 43.2 41.4 20/20 42.3 6.7 min Kimi K2.6
8
ChatGPT Agent (CUA) OpenAI · ChatGPT Agent
42.2 43.8 40.7 20/20 42.2 Kimi K2.6
9
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
33.9 37.3 30.4 20/20 33.9 10.1 min Kimi K2.6
10
MiniMax Agent (CUA) MiniMax · MiniMax Agent
28.0 28.7 27.3 20/20 28.0 Kimi K2.6
IF = instruction following, VQ = visual quality; each judged 0–5 per case by a VLM judge. Scores map to a percentage on a smooth curve where 5/5 = 100% and 4/5 ≈ 92% (a perfect score is the true 100%, with diminishing returns near the top); missing/unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.
Layout, positioning, and spatial-reasoning cases. 10 systems · 7 cases · higher is better

PPTArena score

Combined instruction following + visual quality, normalized to 100

64.3
63.8
63.1
59.8
56.4
48.9
46.0
30.5
26.5
26.3
Claude (CUA)
Codex (GPT-5.5 xhigh)
Gemini CLI (3.5 Flash)
Claude Code (Opus 4.8)
G OpenCode (GLM-5.2)
OpenCode (Kimi K2.7 Code)
D OpenCode (DeepSeek V4 Pro)
ChatGPT Agent (CUA)
OpenCode (MiniMax-M3)
MiniMax Agent (CUA)
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Claude (CUA) Anthropic · Claude 3.7 Sonnet
64.3 66.2 62.5 7/7 64.3 Kimi K2.6
2
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
63.8 56.7 70.9 7/7 63.8 4.8 min Kimi K2.6
3
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
63.1 59.8 66.4 7/7 63.1 4.2 min Kimi K2.6
4
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
59.8 57.5 62.0 7/7 59.8 2.3 min Kimi K2.6
5
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
56.4 56.7 56.1 7/7 56.4 12.0 min Kimi K2.6
6
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
48.9 40.0 57.8 7/7 48.9 12.2 min Kimi K2.6
7
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
46.0 44.2 47.8 7/7 46.0 6.7 min Kimi K2.6
8
ChatGPT Agent (CUA) OpenAI · ChatGPT Agent
30.5 39.4 21.6 7/7 30.5 Kimi K2.6
9
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
26.5 24.2 28.9 7/7 26.5 10.1 min Kimi K2.6
10
MiniMax Agent (CUA) MiniMax · MiniMax Agent
26.3 24.2 28.4 7/7 26.3 Kimi K2.6
IF = instruction following, VQ = visual quality; each judged 0–5 per case by a VLM judge. Scores map to a percentage on a smooth curve where 5/5 = 100% and 4/5 ≈ 92% (a perfect score is the true 100%, with diminishing returns near the top); missing/unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.
Visual styling, theme, and typography-sensitive cases. 10 systems · 6 cases · higher is better

PPTArena score

Combined instruction following + visual quality, normalized to 100

53.0
53.0
49.9
48.6
42.4
36.2
35.0
34.4
30.6
25.4
Codex (GPT-5.5 xhigh)
Claude Code (Opus 4.8)
G OpenCode (GLM-5.2)
Gemini CLI (3.5 Flash)
Claude (CUA)
D OpenCode (DeepSeek V4 Pro)
ChatGPT Agent (CUA)
OpenCode (MiniMax-M3)
OpenCode (Kimi K2.7 Code)
MiniMax Agent (CUA)
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
53.0 56.3 49.7 6/6 53.0 4.8 min Kimi K2.6
2
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
53.0 60.5 45.4 6/6 53.0 2.3 min Kimi K2.6
3
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
49.9 58.2 41.6 6/6 49.9 12.0 min Kimi K2.6
4
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
48.6 57.9 39.3 6/6 48.6 4.2 min Kimi K2.6
5
Claude (CUA) Anthropic · Claude 3.7 Sonnet
42.4 54.6 30.1 6/6 42.4 Kimi K2.6
6
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
36.2 39.3 33.2 6/6 36.2 6.7 min Kimi K2.6
7
ChatGPT Agent (CUA) OpenAI · ChatGPT Agent
35.0 28.2 41.7 6/6 35.0 Kimi K2.6
8
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
34.4 34.4 34.4 6/6 34.4 10.1 min Kimi K2.6
9
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
30.6 36.0 25.2 6/6 30.6 12.2 min Kimi K2.6
10
MiniMax Agent (CUA) MiniMax · MiniMax Agent
25.4 22.6 28.2 6/6 25.4 Kimi K2.6
IF = instruction following, VQ = visual quality; each judged 0–5 per case by a VLM judge. Scores map to a percentage on a smooth curve where 5/5 = 100% and 4/5 ≈ 92% (a perfect score is the true 100%, with diminishing returns near the top); missing/unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.
Structural, cross-slide, and document-level edits. 10 systems · 3 cases · higher is better

PPTArena score

Combined instruction following + visual quality, normalized to 100

84.6
72.0
72.0
65.7
55.3
49.8
45.6
41.8
40.0
30.8
Claude (CUA)
Claude Code (Opus 4.8)
Gemini CLI (3.5 Flash)
ChatGPT Agent (CUA)
G OpenCode (GLM-5.2)
D OpenCode (DeepSeek V4 Pro)
MiniMax Agent (CUA)
Codex (GPT-5.5 xhigh)
OpenCode (Kimi K2.7 Code)
OpenCode (MiniMax-M3)
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Claude (CUA) Anthropic · Claude 3.7 Sonnet
84.6 89.4 79.7 3/3 84.6 Kimi K2.6
2
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
72.0 61.6 82.4 3/3 72.0 2.3 min Kimi K2.6
3
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
72.0 76.5 67.4 3/3 72.0 4.2 min Kimi K2.6
4
ChatGPT Agent (CUA) OpenAI · ChatGPT Agent
65.7 82.4 49.0 3/3 65.7 Kimi K2.6
5
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
55.3 51.7 58.9 3/3 55.3 12.0 min Kimi K2.6
6
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
49.8 53.0 46.6 3/3 49.8 6.7 min Kimi K2.6
7
MiniMax Agent (CUA) MiniMax · MiniMax Agent
45.6 53.0 38.1 3/3 45.6 Kimi K2.6
8
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
41.8 43.2 40.5 3/3 41.8 4.8 min Kimi K2.6
9
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
40.0 43.2 36.8 3/3 40.0 12.2 min Kimi K2.6
10
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
30.8 33.3 28.2 3/3 30.8 10.1 min Kimi K2.6
IF = instruction following, VQ = visual quality; each judged 0–5 per case by a VLM judge. Scores map to a percentage on a smooth curve where 5/5 = 100% and 4/5 ≈ 92% (a perfect score is the true 100%, with diminishing returns near the top); missing/unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.
All 100 PPTArena cases, with unscored cases counted as 0. 7 systems · 100 cases · higher is better

PPTArena score

Combined instruction following + visual quality, normalized to 100

68.0
64.8
55.8
48.7
42.8
40.4
33.9
Codex (GPT-5.5 xhigh)
Claude Code (Opus 4.8)
Gemini CLI (3.5 Flash)
OpenCode (Kimi K2.7 Code)
D OpenCode (DeepSeek V4 Pro)
G OpenCode (GLM-5.2)
OpenCode (MiniMax-M3)
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
68.0 65.4 70.6 100/100 68.0 4.8 min Kimi K2.6
2
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
64.8 64.6 65.0 100/100 64.8 2.3 min Kimi K2.6
3
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
55.8 63.7 47.9 100/100 55.8 4.2 min Kimi K2.6
4
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
48.7 48.0 49.3 100/100 48.7 12.2 min Kimi K2.6
5
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
42.8 48.0 37.6 100/100 42.8 6.7 min Kimi K2.6
6
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
40.4 47.4 33.5 100/100 40.4 12.0 min Kimi K2.6
7
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
33.9 38.8 29.1 100/100 33.9 10.1 min Kimi K2.6
IF = instruction following, VQ = visual quality; each judged 0–5 per case by a VLM judge. Scores map to a percentage on a smooth curve where 5/5 = 100% and 4/5 ≈ 92% (a perfect score is the true 100%, with diminishing returns near the top); missing/unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.
Cases tagged for content and semantic editing. 7 systems · 67 cases · higher is better

PPTArena score

Combined instruction following + visual quality, normalized to 100

69.1
64.8
56.8
51.4
44.6
43.9
34.7
Codex (GPT-5.5 xhigh)
Claude Code (Opus 4.8)
Gemini CLI (3.5 Flash)
OpenCode (Kimi K2.7 Code)
D OpenCode (DeepSeek V4 Pro)
G OpenCode (GLM-5.2)
OpenCode (MiniMax-M3)
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
69.1 67.0 71.1 67/67 69.1 4.8 min Kimi K2.6
2
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
64.8 65.4 64.2 67/67 64.8 2.3 min Kimi K2.6
3
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
56.8 64.9 48.6 67/67 56.8 4.2 min Kimi K2.6
4
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
51.4 52.5 50.3 67/67 51.4 12.2 min Kimi K2.6
5
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
44.6 50.7 38.6 67/67 44.6 6.7 min Kimi K2.6
6
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
43.9 53.3 34.4 67/67 43.9 12.0 min Kimi K2.6
7
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
34.7 40.3 29.2 67/67 34.7 10.1 min Kimi K2.6
IF = instruction following, VQ = visual quality; each judged 0–5 per case by a VLM judge. Scores map to a percentage on a smooth curve where 5/5 = 100% and 4/5 ≈ 92% (a perfect score is the true 100%, with diminishing returns near the top); missing/unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.
Layout, positioning, and spatial-reasoning cases. 7 systems · 29 cases · higher is better

PPTArena score

Combined instruction following + visual quality, normalized to 100

71.3
67.4
59.7
48.4
47.1
35.0
34.8
Codex (GPT-5.5 xhigh)
Claude Code (Opus 4.8)
Gemini CLI (3.5 Flash)
D OpenCode (DeepSeek V4 Pro)
OpenCode (Kimi K2.7 Code)
G OpenCode (GLM-5.2)
OpenCode (MiniMax-M3)
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
71.3 66.3 76.3 29/29 71.3 4.8 min Kimi K2.6
2
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
67.4 64.3 70.4 29/29 67.4 2.3 min Kimi K2.6
3
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
59.7 62.2 57.2 29/29 59.7 4.2 min Kimi K2.6
4
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
48.4 51.1 45.8 29/29 48.4 6.7 min Kimi K2.6
5
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
47.1 40.7 53.4 29/29 47.1 12.2 min Kimi K2.6
6
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
35.0 36.0 34.0 29/29 35.0 12.0 min Kimi K2.6
7
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
34.8 36.1 33.4 29/29 34.8 10.1 min Kimi K2.6
IF = instruction following, VQ = visual quality; each judged 0–5 per case by a VLM judge. Scores map to a percentage on a smooth curve where 5/5 = 100% and 4/5 ≈ 92% (a perfect score is the true 100%, with diminishing returns near the top); missing/unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.
Visual styling, theme, and typography-sensitive cases. 7 systems · 29 cases · higher is better

PPTArena score

Combined instruction following + visual quality, normalized to 100

64.1
63.3
53.8
45.7
39.5
39.0
26.5
Claude Code (Opus 4.8)
Codex (GPT-5.5 xhigh)
Gemini CLI (3.5 Flash)
OpenCode (Kimi K2.7 Code)
D OpenCode (DeepSeek V4 Pro)
G OpenCode (GLM-5.2)
OpenCode (MiniMax-M3)
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
64.1 65.1 63.0 29/29 64.1 2.3 min Kimi K2.6
2
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
63.3 62.2 64.4 29/29 63.3 4.8 min Kimi K2.6
3
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
53.8 63.8 43.8 29/29 53.8 4.2 min Kimi K2.6
4
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
45.7 48.5 42.9 29/29 45.7 12.2 min Kimi K2.6
5
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
39.5 48.0 31.0 29/29 39.5 6.7 min Kimi K2.6
6
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
39.0 50.7 27.2 29/29 39.0 12.0 min Kimi K2.6
7
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
26.5 34.5 18.6 29/29 26.5 10.1 min Kimi K2.6
IF = instruction following, VQ = visual quality; each judged 0–5 per case by a VLM judge. Scores map to a percentage on a smooth curve where 5/5 = 100% and 4/5 ≈ 92% (a perfect score is the true 100%, with diminishing returns near the top); missing/unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.
Structural, cross-slide, and document-level edits. 7 systems · 15 cases · higher is better

PPTArena score

Combined instruction following + visual quality, normalized to 100

64.2
62.4
51.6
47.3
36.7
29.8
27.0
Codex (GPT-5.5 xhigh)
Claude Code (Opus 4.8)
Gemini CLI (3.5 Flash)
OpenCode (Kimi K2.7 Code)
D OpenCode (DeepSeek V4 Pro)
OpenCode (MiniMax-M3)
G OpenCode (GLM-5.2)
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
64.2 62.5 65.8 15/15 64.2 4.8 min Kimi K2.6
2
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
62.4 59.6 65.1 15/15 62.4 2.3 min Kimi K2.6
3
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
51.6 60.8 42.3 15/15 51.6 4.2 min Kimi K2.6
4
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
47.3 44.3 50.3 15/15 47.3 12.2 min Kimi K2.6
5
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
36.7 40.7 32.8 15/15 36.7 6.7 min Kimi K2.6
6
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
29.8 35.8 23.7 15/15 29.8 10.1 min Kimi K2.6
7
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
27.0 32.8 21.1 15/15 27.0 12.0 min Kimi K2.6
IF = instruction following, VQ = visual quality; each judged 0–5 per case by a VLM judge. Scores map to a percentage on a smooth curve where 5/5 = 100% and 4/5 ≈ 92% (a perfect score is the true 100%, with diminishing returns near the top); missing/unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.
Transitions, animations, actions, and interactive features. 7 systems · 4 cases · higher is better

PPTArena score

Combined instruction following + visual quality, normalized to 100

89.4
89.4
64.5
44.4
35.6
35.6
35.6
Codex (GPT-5.5 xhigh)
Claude Code (Opus 4.8)
Gemini CLI (3.5 Flash)
OpenCode (Kimi K2.7 Code)
G OpenCode (GLM-5.2)
OpenCode (MiniMax-M3)
D OpenCode (DeepSeek V4 Pro)
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
89.4 94.1 84.8 4/4 89.4 4.8 min Kimi K2.6
2
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
89.4 94.1 84.8 4/4 89.4 2.3 min Kimi K2.6
3
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
64.5 94.1 35.0 4/4 64.5 4.2 min Kimi K2.6
4
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
44.4 50.0 38.8 4/4 44.4 12.2 min Kimi K2.6
5
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
35.6 50.0 21.2 4/4 35.6 12.0 min Kimi K2.6
6
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
35.6 50.0 21.2 4/4 35.6 10.1 min Kimi K2.6
7
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
35.6 50.0 21.2 4/4 35.6 6.7 min Kimi K2.6
IF = instruction following, VQ = visual quality; each judged 0–5 per case by a VLM judge. Scores map to a percentage on a smooth curve where 5/5 = 100% and 4/5 ≈ 92% (a perfect score is the true 100%, with diminishing returns near the top); missing/unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.

Methodology

What PPTArena measures and how systems are scored.

Task suite

100 real-world decks with 2,125 slides and 1,300+ paired edit instructions. Every case ships the original deck, a natural-language instruction, and a human-made ground truth, spanning five edit categories.

Content 67 Layout 29 Styling 29 Structure 15 Interactivity 4

Scoring

A VLM judge compares each prediction against the ground truth and scores instruction following and visual quality from 0–5. Each is mapped to a percentage on a smooth curve (5/5 = 100%, 4/5 ≈ 92%, 3/5 ≈ 76%) that rewards a perfect result while keeping diminishing returns near the top. A system's split score is the mean over all expected cases, with missing or unscored cases counted as 0.

Splits & judges

The full set covers all 100 cases; the hard subset is a matched 25-case slice used for cost-sensitive agent comparisons. Runs are judged with GPT-5.2 or Gemini 3.1 Pro, and every case can be re-run and re-judged live below.

Live evaluation

Pick any benchmark case, generate a prediction with your own API keys, and score it with the same VLM judge used for the leaderboard.

Case: Case 1: Translate to English But Keep French

1 Select test case

Prompt Please translate this presentation from Kazakh to English. However, it's very important that you do not translate any of the French text. Leave all French words and sentences exactly as they are.
Style target Translate all Kazakh text to English, while preserving all French text without any changes. The specific Kazakh text to be translated is as follows: - Slide 4: The text after the equals sign. - Slide 5: The first paragraph. - Slide 6: The numbered list. - Slides 8 & 9: The text in the right column of each table. - Slide 10: The text in the right-hand text box. - Slide 12: The text in the right-hand text box and the Kazakh word in the title. - Slide 17: The title and the main text box. All original text formatting (font, size, style, color) must be preserved for both the translated English text and the untouched French text. The overall layout and structure of each slide should be preserved, with elements remaining in their approximate original positions. All content must be readable and free of overlaps. No elements should be added or deleted.

2 Generate prediction

Used when running Loop for iterative refinement.

3 LLM judge

The judge compares ground truth vs prediction, with the original deck as baseline.

Status & progress
Live poller stops automatically after ~2 minutes.

Ground truth

TranslatetoEnglishButKeepallFrenchGroundTruthA.pptx

Original

TranslateToEnglishButKeepallFrenchTestA_Bonjour mes amis!.pptx