Technical report / September 2026

GPT-6-Astra
in a Navigation Workflow

Behavioral Analysis in Zero-Shot Vision-and-Language
Navigation in Continuous Environments

Guangzhao Dai1·Qi Wu2·Bin Zhu1,†

1 School of Computing and Information Systems, Singapore Management University

2 Australia Institute for Machine Learning

† Corresponding author and project lead

Illustrated navigation journey connecting visual-language grounding, spatial judgments, temporal reasoning, action adjustment, and evidence-guided verification within an externally managed workflow.
Five perspectives on one navigation task. The mountain is a visual metaphor, not a capability hierarchy.
50 episodes

9 scenes · R2R-CE val-unseen

52.0%

Success rate

48.9%

SPL

70.8%

nDTW

Complete-system results on 50 episodes selected from the 100-episode evaluation pool used by Open-Nav.

Three findings

Correct judgments.
What happens next?

We follow model responses through observations, executed actions, and the end of each trajectory.

01

GPT-6-Astra shows strengths in linking landmarks and earlier actions to instructions.

Recorded responses distinguish the “second on the left” from neighboring openings and use later views to confirm an earlier turn’s instruction match. These judgments use observations and workflow-supplied history.

Follow the evidence
02

GPT-6-Astra shows strengths in seeking visual information and revising uncertain judgments.

Additional views reveal a hidden passage or support rejecting an uncertain hallway candidate. Across the run, 36 of 66 adjustment processes show local relief, and 48 of 75 verification processes clarify a judgment with supporting evidence.

See an adjustment
03

GPT-6-Astra shows limitations in translating task understanding into autonomous completion within this workflow.

An unfinished crossing is recognized while rotation continues. Although 52.0% of episodes meet the endpoint success criterion, only 36.0% both succeed and end with a workflow-accepted STOP; another 16.0% succeed at the step limit.

Inspect the outcomes

Recorded navigation

Watch the route unfold.

Four complete observation replays from the evaluated run. Follow a useful adjustment, a successful stop, or the moment progress stalls.

EP244 · original RGB observationsStep 0 / 45
23 s · 46 observationsDownload replay ↓

Action adjustment

Moving past the door reveals the passage.

Instruction excerptGo through the door and turn left.

At Step 23, GPT-6-Astra proposes moving forward to see past a door leaf. The next observation exposes the left-side floor; subsequent turns align the view with the passage.

Success · workflow-accepted STOPFinal navigation error: 1.56 m · 45 actions

Jump to an observation

Original saved RGB frames, shown at 2 observations per second. Step t is the observation after t actions; model-call delays are omitted. Descriptions also use the saved decisions and execution receipts. These selected cases illustrate behavior, not its frequency.

The study

Abstract

We study GPT-6-Astra in a zero-shot Vision-and-Language Navigation in Continuous Environments (VLN-CE) system, where it interprets instructions, assesses its surroundings, and proposes actions. The system uses a common observation–decision–execution workflow with direct model API calls, without a packaged agent harness or navigation-specific fine-tuning. In this workflow, each request receives selected observations, execution feedback, and retained progress records. Evaluation covers the complete system, including context management and action control.

We evaluate the system on 50 of the 100 R2R-CE val-unseen episodes used by Open-Nav. It achieves a success rate of 52.0%, an SPL of 48.9%, and an nDTW of 70.8%. Our analysis highlights three findings. First, recorded responses link landmarks and earlier actions to instructions using observations and supplied history. Second, reviews include requests for additional views and revisions of uncertain judgments. Third, the results suggest a gap between task understanding and autonomous completion: an unfinished crossing is recognized while rotation continues.

At termination, 36.0% of episodes succeed with a workflow-accepted STOP, while another 16.0% meet the distance criterion at the step limit. These results highlight a central challenge: translating correct local judgments into sustained progress and appropriate stopping.

The navigation workflow

Observe. Decide. Execute.

Direct model API calls, without a packaged agent harness or navigation-specific fine-tuning.

Workflow showing observations, externally retained state, four GPT-6-Astra call roles, action constraints, and execution feedback.
Figure 2. Model calls within an externally managed navigation loop. View PDF ↗

GPT-6-Astra

Interprets the instruction, assesses progress, proposes actions, and reviews arrival evidence.

External workflow

Selects context, retains records, schedules reviews, and applies action and stopping constraints.

Evaluation

Links model judgments to actual motion. Results describe the complete system under these conditions.

Inside the trajectories

Five perspectives.
Evidence at every step.

Explore the analyses behind the three findings. Each perspective connects an interpretation to what the system actually did.

Visual–language grounding

A category is only the beginning.

Identifying a bathroom does not establish which bathroom the instruction means. Qualifiers such as “second on the left” must agree with the observed route.

33 / 50
Category supported
23 / 50
Relation supported
18 / 50
Target instance supported

One terminal destination reference per episode. These counts measure supported evidence; missing or ambiguous evidence is not automatically an error.

Read Section 5.1 ↗
Nine grounding cases connecting instruction phrases, original visual observations, and executed trajectories.
Figure 4 · Full-resolution PDF ↗

Spatial judgments and executed motion

Seeing an entrance is not crossing it.

Position, heading, passage clearance, and route connectivity determine what a visible opening means for the next action.

EP116 · a stalled crossing

The bedroom is visible, but recorded restrictions block forward movement. Correctly identifying the pending crossing does not make the crossing happen.

Cases combine model judgments, controller constraints, and execution receipts. They do not establish that bypassing a restriction would be safe.

Read Section 5.2 ↗
Six navigation examples pairing RGB observations with heading, local geometry, movement restrictions, and route connectivity.
Figure 5 · Full-resolution PDF ↗

Temporal reasoning with supplied history

An action can happen before its meaning becomes clear.

In EP423, later views confirm that an earlier turn matched the instruction. The time of the action remains distinct from the time of confirmation.

31
All timing checks supported
31
All order checks supported
14
All continuity checks supported

Counts are episodes with support for every applicable check. Judgments use workflow-supplied history; different eligibility rules prevent ranking these as independent abilities.

Read Section 5.3 ↗
Stacked bars for event timing, order constraints, and continuity, showing fully supported evidence in 31, 31, and 14 episodes respectively.
Figure 6 · Full-resolution PDF ↗

Action adjustment and local outcomes

A small move can reveal what repeated turns cannot.

In EP244, GPT-6-Astra proposes moving past a nearby door leaf. The executed movement reveals the hidden continuation.

36 / 66

adjustment processes show local relief

42 episodes contain eligible processes. Local relief is distinct from final navigation success.

Read Section 5.4 ↗
Strategy shares and local outcome bars, alongside EP244's move past a door leaf and subsequent view of a passage.
Figure 7 · Full-resolution PDF ↗

Evidence-guided verification

New evidence can rule out a plausible candidate.

In EP308, additional views reveal shelving and a back wall. GPT-6-Astra rejects the uncertain hallway candidate, while the correct route remains unresolved.

48 / 75

verification processes clarify an uncertain judgment

Reviews are workflow-scheduled. The inventory does not measure an error-correction rate.

Read Section 5.5 ↗
Verification outcomes and episode coverage, with sequential views showing why EP308 rejects a storage opening as the instructed hallway.
Figure 8 · Full-resolution PDF ↗

Arrival and stopping

Reaching the goal.
Knowing when to stop.

Endpoint success and accepted stopping tell different parts of the story.

52%

Endpoint success

36%

Success + accepted STOP

16%

Success at step limit

  • Success · accepted STOP
  • Success · step limit
  • Failure · accepted STOP
  • Failure · step limit

Counts cover all 50 episodes. Success uses a 3 m endpoint distance criterion; accepted STOP combines model assessment with workflow checks.

View the complete outcome table
Endpoint success and workflow-accepted stopping
TerminationSuccessFailureTotal
Workflow-accepted STOP18321
Budget termination82129
Total262450

Interpreting the findings

What this study
establishes.

These findings characterize GPT-6-Astra within the evaluated workflow and subset. Context selection, retained records, review scheduling, and action control all shape the observed behavior.

One run on a fixed 50-episode subset does not isolate the model’s contribution or establish full-benchmark superiority. The subset was not held out from all workflow development, and the retrospective, assistant-assisted annotations have not undergone independent human validation.

Discussion and limitations ↗
Figure preview