P PPTArena
API keys

Keys stay in your browser's local storage and are only sent with requests you run.

The first benchmark for editing real PowerPoint decks.

Michael Ofengenden · Yunze Man · Ziqi Pang · Liang-Yan Gui · Yu-Xiong Wang

100 Real decks
2,125 Slides
1,300+ Paired edits
IF + VQ Judge metrics

Leaderboard

Matched 25-case hard subset used for cost-sensitive agent comparisons. 11 systems · 25 cases · higher is better

PPTArena score

Mean instruction following + visual quality, shown as a linear percentage of available points

55.6
48.4
43.6
42.8
42.4
41.6
36.0
33.6
33.2
26.0
21.6
DuMate
Claude CUA
GPT-5.5
Claude 4.8
Gemini 3.5
G GLM-5.2
Kimi K2.7
ChatGPT CUA
D DeepSeek V4
MiniMax-M3
MiniMax CUA
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Baidu DuMate Baidu · PPTX Skill r266 · Moonshot native
55.6 48.8 62.4 25/25 55.6 — Kimi K2.6
2
Claude (CUA) Anthropic · Claude 3.7 Sonnet
48.4 55.2 41.6 25/25 48.4 — Kimi K2.6
3
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
43.6 43.2 44.0 25/25 43.6 4.8 min Kimi K2.6
4
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
42.8 44.8 40.8 25/25 42.8 2.3 min Kimi K2.6
5
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
42.4 48.0 36.8 25/25 42.4 4.2 min Kimi K2.6
6
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
41.6 46.4 36.8 25/25 41.6 12.0 min Kimi K2.6
7
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
36.0 35.2 36.8 25/25 36.0 12.2 min Kimi K2.6
8
ChatGPT Agent (CUA) OpenAI · ChatGPT Agent
33.6 36.0 31.2 25/25 33.6 — Kimi K2.6
9
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
33.2 34.4 32.0 25/25 33.2 6.7 min Kimi K2.6
10
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
26.0 28.0 24.0 25/25 26.0 10.1 min Kimi K2.6
11
MiniMax Agent (CUA) MiniMax · MiniMax Agent
21.6 23.2 20.0 25/25 21.6 — Kimi K2.6
IF = instruction following, VQ = visual quality; each is judged 0–5 per case by a VLM judge and linearly scaled to a percentage of available points. Missing or unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.
Cases tagged for content and semantic editing. 11 systems · 20 cases · higher is better

PPTArena score

Mean instruction following + visual quality, shown as a linear percentage of available points

52.5
47.5
44.5
41.0
39.5
39.0
37.0
32.5
32.0
26.0
20.0
DuMate
Claude CUA
GPT-5.5
G GLM-5.2
Claude 4.8
Gemini 3.5
Kimi K2.7
D DeepSeek V4
ChatGPT CUA
MiniMax-M3
MiniMax CUA
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Baidu DuMate Baidu · PPTX Skill r266 · Moonshot native
52.5 47.0 58.0 20/20 52.5 — Kimi K2.6
2
Claude (CUA) Anthropic · Claude 3.7 Sonnet
47.5 55.0 40.0 20/20 47.5 — Kimi K2.6
3
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
44.5 45.0 44.0 20/20 44.5 4.8 min Kimi K2.6
4
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
41.0 47.0 35.0 20/20 41.0 12.0 min Kimi K2.6
5
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
39.5 44.0 35.0 20/20 39.5 2.3 min Kimi K2.6
6
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
39.0 45.0 33.0 20/20 39.0 4.2 min Kimi K2.6
7
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
37.0 36.0 38.0 20/20 37.0 12.2 min Kimi K2.6
8
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
32.5 33.0 32.0 20/20 32.5 6.7 min Kimi K2.6
9
ChatGPT Agent (CUA) OpenAI · ChatGPT Agent
32.0 33.0 31.0 20/20 32.0 — Kimi K2.6
10
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
26.0 29.0 23.0 20/20 26.0 10.1 min Kimi K2.6
11
MiniMax Agent (CUA) MiniMax · MiniMax Agent
20.0 21.0 19.0 20/20 20.0 — Kimi K2.6
IF = instruction following, VQ = visual quality; each is judged 0–5 per case by a VLM judge and linearly scaled to a percentage of available points. Missing or unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.
Layout, positioning, and spatial-reasoning cases. 11 systems · 7 cases · higher is better

PPTArena score

Mean instruction following + visual quality, shown as a linear percentage of available points

52.9
51.4
50.0
48.6
45.7
42.9
37.1
32.9
22.9
20.0
18.6
Claude CUA
DuMate
GPT-5.5
Gemini 3.5
Claude 4.8
G GLM-5.2
Kimi K2.7
D DeepSeek V4
ChatGPT CUA
MiniMax-M3
MiniMax CUA
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Claude (CUA) Anthropic · Claude 3.7 Sonnet
52.9 54.3 51.4 7/7 52.9 — Kimi K2.6
2
Baidu DuMate Baidu · PPTX Skill r266 · Moonshot native
51.4 48.6 54.3 7/7 51.4 — Kimi K2.6
3
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
50.0 42.9 57.1 7/7 50.0 4.8 min Kimi K2.6
4
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
48.6 45.7 51.4 7/7 48.6 4.2 min Kimi K2.6
5
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
45.7 42.9 48.6 7/7 45.7 2.3 min Kimi K2.6
6
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
42.9 42.9 42.9 7/7 42.9 12.0 min Kimi K2.6
7
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
37.1 28.6 45.7 7/7 37.1 12.2 min Kimi K2.6
8
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
32.9 31.4 34.3 7/7 32.9 6.7 min Kimi K2.6
9
ChatGPT Agent (CUA) OpenAI · ChatGPT Agent
22.9 28.6 17.1 7/7 22.9 — Kimi K2.6
10
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
20.0 17.1 22.9 7/7 20.0 10.1 min Kimi K2.6
11
MiniMax Agent (CUA) MiniMax · MiniMax Agent
18.6 17.1 20.0 7/7 18.6 — Kimi K2.6
IF = instruction following, VQ = visual quality; each is judged 0–5 per case by a VLM judge and linearly scaled to a percentage of available points. Missing or unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.
Visual styling, theme, and typography-sensitive cases. 11 systems · 6 cases · higher is better

PPTArena score

Mean instruction following + visual quality, shown as a linear percentage of available points

56.7
41.7
41.7
38.3
36.7
33.3
26.7
26.7
25.0
23.3
18.3
DuMate
GPT-5.5
Claude 4.8
G GLM-5.2
Gemini 3.5
Claude CUA
MiniMax-M3
D DeepSeek V4
ChatGPT CUA
Kimi K2.7
MiniMax CUA
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Baidu DuMate Baidu · PPTX Skill r266 · Moonshot native
56.7 43.3 70.0 6/6 56.7 — Kimi K2.6
2
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
41.7 43.3 40.0 6/6 41.7 4.8 min Kimi K2.6
3
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
41.7 46.7 36.7 6/6 41.7 2.3 min Kimi K2.6
4
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
38.3 46.7 30.0 6/6 38.3 12.0 min Kimi K2.6
5
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
36.7 43.3 30.0 6/6 36.7 4.2 min Kimi K2.6
6
Claude (CUA) Anthropic · Claude 3.7 Sonnet
33.3 43.3 23.3 6/6 33.3 — Kimi K2.6
7
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
26.7 26.7 26.7 6/6 26.7 10.1 min Kimi K2.6
8
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
26.7 30.0 23.3 6/6 26.7 6.7 min Kimi K2.6
9
ChatGPT Agent (CUA) OpenAI · ChatGPT Agent
25.0 20.0 30.0 6/6 25.0 — Kimi K2.6
10
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
23.3 26.7 20.0 6/6 23.3 12.2 min Kimi K2.6
11
MiniMax Agent (CUA) MiniMax · MiniMax Agent
18.3 16.7 20.0 6/6 18.3 — Kimi K2.6
IF = instruction following, VQ = visual quality; each is judged 0–5 per case by a VLM judge and linearly scaled to a percentage of available points. Missing or unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.
Structural, cross-slide, and document-level edits. 11 systems · 3 cases · higher is better

PPTArena score

Mean instruction following + visual quality, shown as a linear percentage of available points

80.0
73.3
63.3
63.3
56.7
46.7
40.0
36.7
36.7
33.3
26.7
DuMate
Claude CUA
Claude 4.8
Gemini 3.5
ChatGPT CUA
G GLM-5.2
D DeepSeek V4
GPT-5.5
MiniMax CUA
Kimi K2.7
MiniMax-M3
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Baidu DuMate Baidu · PPTX Skill r266 · Moonshot native
80.0 66.7 93.3 3/3 80.0 — Kimi K2.6
2
Claude (CUA) Anthropic · Claude 3.7 Sonnet
73.3 80.0 66.7 3/3 73.3 — Kimi K2.6
3
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
63.3 53.3 73.3 3/3 63.3 2.3 min Kimi K2.6
4
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
63.3 73.3 53.3 3/3 63.3 4.2 min Kimi K2.6
5
ChatGPT Agent (CUA) OpenAI · ChatGPT Agent
56.7 73.3 40.0 3/3 56.7 — Kimi K2.6
6
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
46.7 46.7 46.7 3/3 46.7 12.0 min Kimi K2.6
7
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
40.0 46.7 33.3 3/3 40.0 6.7 min Kimi K2.6
8
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
36.7 40.0 33.3 3/3 36.7 4.8 min Kimi K2.6
9
MiniMax Agent (CUA) MiniMax · MiniMax Agent
36.7 46.7 26.7 3/3 36.7 — Kimi K2.6
10
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
33.3 40.0 26.7 3/3 33.3 12.2 min Kimi K2.6
11
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
26.7 33.3 20.0 3/3 26.7 10.1 min Kimi K2.6
IF = instruction following, VQ = visual quality; each is judged 0–5 per case by a VLM judge and linearly scaled to a percentage of available points. Missing or unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.
All 100 PPTArena cases, with unscored cases counted as 0. 8 systems · 100 cases · higher is better

PPTArena score

Mean instruction following + visual quality, shown as a linear percentage of available points

58.1
57.6
54.2
45.7
41.5
34.9
33.5
28.2
DuMate
GPT-5.5
Claude 4.8
Gemini 3.5
Kimi K2.7
D DeepSeek V4
G GLM-5.2
MiniMax-M3
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Baidu DuMate Baidu · PPTX Skill r266 · Moonshot native
58.1 53.4 62.8 100/100 58.1 — Kimi K2.6
2
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
57.6 54.4 60.8 100/100 57.6 4.8 min Kimi K2.6
3
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
54.2 53.2 55.2 100/100 54.2 2.3 min Kimi K2.6
4
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
45.7 53.0 38.4 100/100 45.7 4.2 min Kimi K2.6
5
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
41.5 40.4 42.6 100/100 41.5 12.2 min Kimi K2.6
6
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
34.9 40.2 29.6 100/100 34.9 6.7 min Kimi K2.6
7
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
33.5 40.0 27.0 100/100 33.5 12.0 min Kimi K2.6
8
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
28.2 32.8 23.6 100/100 28.2 10.1 min Kimi K2.6
IF = instruction following, VQ = visual quality; each is judged 0–5 per case by a VLM judge and linearly scaled to a percentage of available points. Missing or unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.
Cases tagged for content and semantic editing. 8 systems · 67 cases · higher is better

PPTArena score

Mean instruction following + visual quality, shown as a linear percentage of available points

58.7
58.4
54.5
46.6
44.0
36.6
36.3
28.8
GPT-5.5
DuMate
Claude 4.8
Gemini 3.5
Kimi K2.7
D DeepSeek V4
G GLM-5.2
MiniMax-M3
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
58.7 55.8 61.5 67/67 58.7 4.8 min Kimi K2.6
2
Baidu DuMate Baidu · PPTX Skill r266 · Moonshot native
58.4 53.1 63.6 67/67 58.4 — Kimi K2.6
3
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
54.5 54.0 54.9 67/67 54.5 2.3 min Kimi K2.6
4
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
46.6 53.7 39.4 67/67 46.6 4.2 min Kimi K2.6
5
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
44.0 44.2 43.9 67/67 44.0 12.2 min Kimi K2.6
6
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
36.6 42.7 30.4 67/67 36.6 6.7 min Kimi K2.6
7
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
36.3 44.8 27.8 67/67 36.3 12.0 min Kimi K2.6
8
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
28.8 33.7 23.9 67/67 28.8 10.1 min Kimi K2.6
IF = instruction following, VQ = visual quality; each is judged 0–5 per case by a VLM judge and linearly scaled to a percentage of available points. Missing or unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.
Layout, positioning, and spatial-reasoning cases. 8 systems · 29 cases · higher is better

PPTArena score

Mean instruction following + visual quality, shown as a linear percentage of available points

59.7
55.9
55.5
49.0
39.7
39.3
29.3
28.6
GPT-5.5
DuMate
Claude 4.8
Gemini 3.5
Kimi K2.7
D DeepSeek V4
G GLM-5.2
MiniMax-M3
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
59.7 53.8 65.5 29/29 59.7 4.8 min Kimi K2.6
2
Baidu DuMate Baidu · PPTX Skill r266 · Moonshot native
55.9 51.7 60.0 29/29 55.9 — Kimi K2.6
3
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
55.5 51.7 59.3 29/29 55.5 2.3 min Kimi K2.6
4
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
49.0 51.0 46.9 29/29 49.0 4.2 min Kimi K2.6
5
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
39.7 33.1 46.2 29/29 39.7 12.2 min Kimi K2.6
6
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
39.3 42.1 36.6 29/29 39.3 6.7 min Kimi K2.6
7
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
29.3 29.7 29.0 29/29 29.3 12.0 min Kimi K2.6
8
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
28.6 29.7 27.6 29/29 28.6 10.1 min Kimi K2.6
IF = instruction following, VQ = visual quality; each is judged 0–5 per case by a VLM judge and linearly scaled to a percentage of available points. Missing or unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.
Visual styling, theme, and typography-sensitive cases. 8 systems · 29 cases · higher is better

PPTArena score

Mean instruction following + visual quality, shown as a linear percentage of available points

53.1
52.4
52.1
42.4
38.6
31.7
31.7
22.1
Claude 4.8
DuMate
GPT-5.5
Gemini 3.5
Kimi K2.7
G GLM-5.2
D DeepSeek V4
MiniMax-M3
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
53.1 53.1 53.1 29/29 53.1 2.3 min Kimi K2.6
2
Baidu DuMate Baidu · PPTX Skill r266 · Moonshot native
52.4 49.0 55.9 29/29 52.4 — Kimi K2.6
3
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
52.1 50.3 53.8 29/29 52.1 4.8 min Kimi K2.6
4
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
42.4 51.0 33.8 29/29 42.4 4.2 min Kimi K2.6
5
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
38.6 40.7 36.6 29/29 38.6 12.2 min Kimi K2.6
6
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
31.7 42.8 20.7 29/29 31.7 12.0 min Kimi K2.6
7
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
31.7 40.0 23.4 29/29 31.7 6.7 min Kimi K2.6
8
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
22.1 29.0 15.2 29/29 22.1 10.1 min Kimi K2.6
IF = instruction following, VQ = visual quality; each is judged 0–5 per case by a VLM judge and linearly scaled to a percentage of available points. Missing or unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.
Structural, cross-slide, and document-level edits. 8 systems · 15 cases · higher is better

PPTArena score

Mean instruction following + visual quality, shown as a linear percentage of available points

60.0
52.7
50.7
42.0
39.3
28.7
24.0
22.0
DuMate
GPT-5.5
Claude 4.8
Gemini 3.5
Kimi K2.7
D DeepSeek V4
MiniMax-M3
G GLM-5.2
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Baidu DuMate Baidu · PPTX Skill r266 · Moonshot native
60.0 52.0 68.0 15/15 60.0 — Kimi K2.6
2
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
52.7 50.7 54.7 15/15 52.7 4.8 min Kimi K2.6
3
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
50.7 48.0 53.3 15/15 50.7 2.3 min Kimi K2.6
4
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
42.0 52.0 32.0 15/15 42.0 4.2 min Kimi K2.6
5
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
39.3 37.3 41.3 15/15 39.3 12.2 min Kimi K2.6
6
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
28.7 33.3 24.0 15/15 28.7 6.7 min Kimi K2.6
7
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
24.0 30.7 17.3 15/15 24.0 10.1 min Kimi K2.6
8
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
22.0 28.0 16.0 15/15 22.0 12.0 min Kimi K2.6
IF = instruction following, VQ = visual quality; each is judged 0–5 per case by a VLM judge and linearly scaled to a percentage of available points. Missing or unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.
Transitions, animations, actions, and interactive features. 8 systems · 4 cases · higher is better

PPTArena score

Mean instruction following + visual quality, shown as a linear percentage of available points

85.0
82.5
82.5
57.5
42.5
32.5
32.5
32.5
DuMate
GPT-5.5
Claude 4.8
Gemini 3.5
Kimi K2.7
G GLM-5.2
MiniMax-M3
D DeepSeek V4
# Model Score IF VQ Cases Avg scored Avg edit Judge
1
Baidu DuMate Baidu · PPTX Skill r266 · Moonshot native
85.0 95.0 75.0 4/4 85.0 — Kimi K2.6
2
Codex (GPT-5.5 xhigh) OpenAI · GPT-5.5 xhigh
82.5 90.0 75.0 4/4 82.5 4.8 min Kimi K2.6
3
Claude Code (Opus 4.8) Anthropic · Claude Opus 4.8
82.5 90.0 75.0 4/4 82.5 2.3 min Kimi K2.6
4
Gemini CLI (3.5 Flash) Google · Gemini 3.5 Flash
57.5 90.0 25.0 4/4 57.5 4.2 min Kimi K2.6
5
OpenCode (Kimi K2.7 Code) Moonshot AI · OpenCode · Kimi K2.7 Code
42.5 50.0 35.0 4/4 42.5 12.2 min Kimi K2.6
6
G
OpenCode (GLM-5.2) Zhipu AI · OpenCode · GLM-5.2
32.5 50.0 15.0 4/4 32.5 12.0 min Kimi K2.6
7
OpenCode (MiniMax-M3) MiniMax · OpenCode · MiniMax-M3
32.5 50.0 15.0 4/4 32.5 10.1 min Kimi K2.6
8
D
OpenCode (DeepSeek V4 Pro) DeepSeek · OpenCode · DeepSeek V4 Pro
32.5 50.0 15.0 4/4 32.5 6.7 min Kimi K2.6
IF = instruction following, VQ = visual quality; each is judged 0–5 per case by a VLM judge and linearly scaled to a percentage of available points. Missing or unscored cases count as 0. Click a column to sort. Entries tagged paper are reported from the PPTArena paper.

Methodology

What PPTArena measures and how systems are scored.

Task suite

100 real-world decks with 2,125 slides and 1,300+ paired edit instructions. Every case ships the original deck, a natural-language instruction, and a human-made ground truth, spanning five edit categories.

Content 67 Layout 29 Styling 29 Structure 15 Interactivity 4

Scoring

A VLM judge compares each prediction against the ground truth and scores instruction following and visual quality from 0–5. Each score is linearly converted to a percentage of available points, so 4/5 = 80% and 5/5 = 100%. A system's split score is the equal-weight mean across both metrics and all expected cases, with missing or unscored cases counted as 0.

Splits & judges

The full set covers all 100 cases; the hard subset is a matched 25-case slice used for cost-sensitive agent comparisons. Each leaderboard row identifies its judge, and every case can be re-run and re-judged live below.

Live evaluation

Pick any benchmark case, generate a prediction with your own API keys, and score it with the same VLM judge used for the leaderboard.

Case: Case 1: Translate to English But Keep French

1 Select test case

Prompt Please translate this presentation from Kazakh to English. However, it's very important that you do not translate any of the French text. Leave all French words and sentences exactly as they are.
Style target Translate all Kazakh text to English, while preserving all French text without any changes. The specific Kazakh text to be translated is as follows: - Slide 4: The text after the equals sign. - Slide 5: The first paragraph. - Slide 6: The numbered list. - Slides 8 & 9: The text in the right column of each table. - Slide 10: The text in the right-hand text box. - Slide 12: The text in the right-hand text box and the Kazakh word in the title. - Slide 17: The title and the main text box. All original text formatting (font, size, style, color) must be preserved for both the translated English text and the untouched French text. The overall layout and structure of each slide should be preserved, with elements remaining in their approximate original positions. All content must be readable and free of overlaps. No elements should be added or deleted.

2 Generate prediction

Used when running Loop for iterative refinement.

3 LLM judge

The judge compares ground truth vs prediction, with the original deck as baseline.

Status & progress
Live poller stops automatically after ~2 minutes.

Ground truth

TranslatetoEnglishButKeepallFrenchGroundTruthA.pptx

Original

TranslateToEnglishButKeepallFrenchTestA_Bonjour mes amis!.pptx