Back
AI agents create editable 3D scenes but struggle to judge accuracy
SiTech AI Team3 min read

AI agents create editable 3D scenes but struggle to judge accuracy

LEGO-Anything turns photos into editable Blender code, while its benchmark finds that coding agents can produce usable scenes but often misjudge geometry and lose accuracy during revision.

From a photo to editable Blender code

LEGO-Anything, a project from researchers at the University of Maryland and AWS, uses a coding agent to turn a single photo into an executable Blender program. The agent writes code, runs it, inspects the result and revises the scene until it resembles the original image.

The approach, called Image-to-Code, produces more than a rendered image. Objects, geometry, layout and camera position are represented explicitly in a program that users can run, inspect, edit and query for image analysis tasks.

A benchmark with simulator ground truth

The team introduced LEGO-Bench to evaluate the resulting scenes. It contains 208 images from 104 indoor and outdoor scenes and uses 443 registered assets. The images are rendered from professionally built simulator scenes, giving the inputs a natural appearance while preserving exact geometry, depth and object assignments as ground truth.

Scene complexity can be increased without changing lighting or camera settings. The benchmark evaluates validity, geometric accuracy and visual similarity to the original. Appearance is measured by re-rendering the submitted scene and comparing it pixel by pixel with the reference image.

Usable scenes, weak geometric judgment

All six tested GPT configurations delivered a working scene almost every time, but their accuracy varied widely. GPT-6 Astra, the best tested model, scored 53.4 percent on indoor scenes and 39.6 percent on outdoor scenes. Some weaker configurations scored around 15 percent. Accuracy declined as scenes became more complex, and outdoor scenes proved harder than interiors.

Increasing the reasoning budget improved the GPT-6 variants. On an office test subset, Astra's score rose from 32.3 percent to 61.8 percent. Analysis of the agents' work identified poor initial attempts, revisions that undid earlier progress and unreliable self-assessment as common problems. The researchers also reported Astra falling from 33.9 percent to 4.4 percent after a late revision.

When asked to choose which of two scene versions better matched the original, the models' geometric judgments were near or below chance. The researchers concluded that refinement should rely on concrete measurements rather than an agent's assessment of its own work.

A plugin and downstream vision tests

The resulting LEGO-Plugin requires no additional training. It anchors the starting scene to the reference image, replaces self-judgment with concrete measurements and protects correct progress from regressive edits. The plugin improved all six tested models, with weaker agents receiving gains of up to 62.7 percent. The strongest model gained about two percentage points.

The team also tested the reconstructed scenes as inputs for object detection, segmentation and depth estimation. Without additional training, they produced usable but unremarkable results. Object detection performed best, reaching roughly half the performance of the specialized DINO model. The gaps with specialized models SAM 3 and Depth Anything 3 were larger for segmentation and depth estimation. The researchers said current coding agents show promise but do not yet create scene programs that are accurate enough for faithful reconstruction.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.