TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 2 October 2026, 01:00 UTC

PPTBench: Can Coding Agents Reconstruct the Visual World through Structured, Editable Slides

What
PPTBench: Can Coding Agents Reconstruct the Visual World through Structured, Editable Slides
Who
arxiv.org
When
1 October 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2609.29718
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

Further analysis shows that increasing reasoning effort primarily improves hard-gate passage rather than mean detail quality on each configuration's passing tasks, while configurations with more inspection tend to achieve higher overall scores. It comes from a paper posted to arXiv on 1 October 2026. Coding agents are increasingly moving beyond text-based software tasks to reconstruct visual targets through code. This capability, visual coding, requires agents to translate their understanding of visual targets into executable code. Slides provide a natural testbed for this capability, combining rich visual structure with objects that can be programmatically created and edited. To measure this capability, we introduce PPTBench, a benchmark for reconstructing scientific flow diagrams as editable PowerPoint slides. PPTBench covers 500 scientific flow-diagram tasks across 10 presentation domains, drawn from real research papers, and is evaluated with a four-stage agentic judge covering artifact validity, process and connector fidelity, rendering quality, and visual fidelity. Evaluation of ten models across 36 model--harness--effort configurations shows a substantial gap in reliable visual coding: the best configuration achieves 77.34, while the median across configurations is 24.38. Fine-grained analysis shows that agents can generally produce valid slide files, but struggle to produce high-quality reconstructions that faithfully recover the semantics and visual structure of the target. PPTBench establishes a measurable testbed for studying visual coding and advancing agents toward more reliable visual creation.

Why it counts

Further analysis shows that increasing reasoning effort primarily improves hard-gate passage rather than mean detail quality on each configuration's passing tasks, while configurations with more inspection tend to achieve higher overall scores.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.

PPTBench: Can Coding Agents Reconstruct the Visual World through Structured, Editable Slides · The Singularity Report