The Human–VLA Communication Gap

What Humans Leave Implicit and What Policies Cannot Recover

Yiqing Xu, Jiajun Wu
Computer Science, Stanford University
Human–VLA communication gap overview
A robot may know how to perform a task yet fail when a human asks for it naturally. For the same physical task, human specifications vary in task content, reference, and composition, producing substantially different VLA performance (top). We separate this mismatch into an information gap, where task semantics, grounding, or event structure remain unresolved, and a residual execution gap, where the fully resolved task remains sensitive to command organization and invocation state (bottom).

Abstract

A robot may know how to perform a task yet fail when a human asks for it naturally. We study this human–VLA communication gap by separating an information gap, where task semantics, grounding, or event structure remain unresolved, from a residual execution gap, where even a fully specified task depends on how it is organized and invoked. To do this, we characterize how people naturally express demonstrated objectives and evaluate controlled variants of the same physical tasks, with detailed analysis on $\pi_{0.5}$ and consistent trends on $\pi_0$-FAST and Cosmos3-Edge. Our study reveals that humans increasingly compress instructions as tasks grow, relying more on contextual grounding and implicit execution structure. Current policies are misaligned with this behavior: they tolerate grounding dependence relatively well, but struggle when task semantics or event structure must be reconstructed. Even after the information gap is resolved, execution remains sensitive to command granularity and invocation state. These findings identify concrete targets for human-centered policy learning and interfaces that better bridge natural human objectives to robot behavior.

Video

Two Gaps Between a Human and a VLA

Ask a robot “Can you clear the table so I can work here?” and it may freeze, even though the same policy succeeds on “Pick up the white mug and place it in the bin.” What failed was not the physical skill but the interface to it. We study this communication gap through two mismatches:

Information gap

A natural specification leaves some of the executable task structure unstated (what to do, on which object, toward what target, as which sequence of events) and the policy must recover it.

Execution gap

Even a complete specification may fail if its resolved events are organized or invoked outside the policy's learned execution interface.

How People Specify Objectives

We asked 50 participants to watch task demonstrations and write the instruction they would give another human, yielding 750 specifications. Informed by these instructions and prior work on situated instruction, we derive a taxonomy that factors any specification by three axes. All examples below describe the same physical task; highlighted spans mark the varied element.

FactorFormExample
Task content Action “Put all the yellow objects in the left bin and all the blue objects in the right bin.”
Outcome “All yellow objects should be in the left bin, and all blue objects should be in the right bin.”
Intent / preference “I want the objects sorted by color.”
Reference form Direct “Put the yellow mug in the left bin and the blue cup in the right bin.”
Grounding-dependent “Put the yellow mug in the bin on the left and the blue cup in the bin on the right.”
Missing “Put the yellow mug and the blue cup away.” (destination omitted)
Composition Enumerated “Put the yellow mug in the left bin, then put the yellow bowl in the left bin, then put the blue cup in the right bin.”
Grouped “Put the yellow mug and bowl in the left bin, and the blue cup and plate in the right bin.”
Collective “Put all the yellow objects in the left bin and all the blue objects in the right bin.”

Measuring the Communication Gap

From each demonstration we recover an ordered sequence of grounded manipulation events $E_\tau=(e_1,\ldots,e_{K_\tau})$, with $e_k=(A_k,O_k,D_k)$ an action, object, and destination. This shared representation lets us compare any human specification $u$ against a fully resolved, policy-facing command built from templates mined from the policy's robot-training language. For each specification we measure:

  • Semantic coverage $\eta_\tau(u)$: the fraction of required action / object / destination components that are stated at all.
  • Groundedness $\gamma_\tau(u)$: the fraction whose grounded values are given directly in language, rather than resolved from scene context.
  • Decomposition $\delta_\tau(u)=J_\tau(u)/K_\tau$: how many of the atomic events are individually expressed.
Quantifying the information gap
The same physical task specified four ways. Less explicit specifications expose less executable structure, and VLA success drops accordingly.

Finding 1

Humans Compress as Tasks Grow

9.1 → 15.8words per instruction
(1 event → 3+ events)
0.89 → 0.27groundedness $\gamma_\tau$
1.00 → 0.21decomposition $\delta_\tau$
Human compression with task complexity

People do not communicate more complex objectives by proportionally lengthening their instructions. Instruction length grows sublinearly with the number of required events, while the executable structure made explicit declines sharply. People preserve much of the task content but increasingly rely on the listener to resolve referents and, especially, to expand a compact expression into atomic manipulations. This compression is structure-dependent: weakly structured tasks stay enumerative, while rule-compressible tasks shift toward outcome-level and collective descriptions.

Finding 2

VLAs Are Selectively Sensitive to What Is Left Implicit

What matters is not simply how much information is unresolved, but what must be recovered. Policies tolerate grounding-dependent references well, but pay much more when task semantics or event structure must be reconstructed. Each undecomposed event costs about as much as a missing action, object, or destination.

Behavioral cost of unresolved task structure
Left: for all three policies, grounding dependence is cheap, while missing components and undecomposed events cost far more. Right: as tasks grow, coverage and decomposition become more strongly associated with the success gap than groundedness.

$\pi_{0.5}$ success across the full factorial space of specifications

ContentReference $\eta_\tau$$\gamma_\tau$ Enumerated
$\bar\delta=1.00$, $N=120$
Grouped
$\bar\delta=0.47$, $N=54$
Collective
$\bar\delta=0.41$, $N=50$
Succ.$\Delta^{\text{info}}$ Succ.$\Delta^{\text{info}}$ Succ.$\Delta^{\text{info}}$

Success (%) and information-gap penalty $\Delta^{\text{info}}_{\text{Success}}$ (points below the fully specified Action–Direct–Enumerated reference). Darker red marks a larger gap; bold marks one-factor ablations from the reference. Means over task-level averages; see the paper for 95% CIs.

Finding 3

A Residual Execution Gap Remains

Is complete executable information sufficient? We fix every compound task to the same fully resolved Action–Direct–Enumerated event sequence and vary only how its events are presented and invoked. No task information is added, yet issuing events one at a time, and resetting the arm to a familiar home pose between them, substantially improves execution for all three policies.

One prompt upfront 23.5%
Sequential, one event at a time 35.6%
Sequential + home reset 47.4%

$\pi_{0.5}$ success on compound tasks with identical grounded semantics.

Residual execution gap across policies
Sequential execution improves success and partial progress over a single upfront prompt for all three policies; a home reset between events adds a further gain.

Example Rollouts

Example rollouts of the information and execution gaps
Left: the policy succeeds when object identity or action is implicit, but fails when task semantics must be inferred. Right: multi-step execution can fail through incorrect object selection; returning to a home pose between subtasks enables completion.

Three Policies, One Gap, Different Weak Points

Across $\pi_{0.5}$, $\pi_0$-FAST, and the world-action model Cosmos3-Edge, the same pattern holds, but each component of the gap traces back to a different part of the model:

  • Grounding is inherited from the vision–language backbone: Cosmos3-Edge pays the least for grounding-dependent references.
  • Organization comes from the robot-language distribution: the sequentialization gain is shared by all DROID-trained policies, and largest for $\pi_0$-FAST (8.2% → 19.5%).
  • Invocation state depends on the action-generation paradigm: Cosmos3-Edge gains 5.8 points from a home reset against 11.8 for $\pi_{0.5}$.
  • Structural reconstruction, the component humans rely on most as tasks grow, is unaddressed by all three, which makes it the primary target for policy learning.

Generalist robots should adapt to how people communicate, rather than requiring people to adapt to the robot's learned interface.