Technical report V2 / September 2026

GPT-6-Astra
Lights Up
Embodied Navigation

Evaluation in Zero-Shot Vision-and-Language
Navigation in Continuous Environments

Guangzhao Dai1Qianru Sun1Qi Wu2Bin Zhu1,†

1 School of Computing and Information Systems, Singapore Management University

2 Australia Institute for Machine Learning

† Corresponding author and project lead

How far can a general-purpose model navigate using its own perception, reasoning, and decisions? Give GPT-6-Astra a monocular camera and basic actions through MCP. It decides when to observe, how to move, and when to stop.

Website & PDF updated 25 September 2026 · arXiv first release: 24 September 2026

Success rate81.3%

Ultra reasoning · monocular RGB

Path efficiency · SPL71.5%

Success weighted by path length

Over zero-shot methods+15.3pp

SR vs. SpatialAnt (66.0%)

Over trained methods+9.2pp

SR vs. Qwen-RobotNav (72.1%)

Comparisons use reported results with different evaluation subsets and system settings. They provide context for the model's performance; they are not a matched-condition ranking.

Two panels compare GPT-6-Astra with zero-shot and trained VLN methods. Medium and ultra reasoning achieve 75.7% and 81.3% success; ultra exceeds the strongest reported results by 15.3 and 9.2 percentage points.
Figure 1 Strong zero-shot navigation performance with monocular RGB. Trained-method results use the full validation-unseen split; GPT-6-Astra uses R2R-CE-100. See scores and evaluation settings ↓
Read the abstract

We investigate whether GPT-6-Astra, a general-purpose foundation model, can navigate unfamiliar environments using its own perception, reasoning, and decision-making capabilities. Our evaluation focuses on zero-shot vision-and-language navigation in continuous environments (VLN-CE) through a minimal interface in the Codex harness, aiming to unleash GPT-6-Astra's full potential for navigation. Using monocular RGB, GPT-6-Astra decides when to observe, how to move, and when to stop, without navigation-specific fine-tuning, a trained waypoint predictor, or a pre-built scene map. Our evaluation yields four key findings and implications. First, GPT-6-Astra achieves strong zero-shot navigation performance using only monocular RGB observations. On R2R-CE-100, GPT-6-Astra (ultra reasoning) achieves a success rate of 81.3%, exceeding the strongest reported zero-shot and even train-based success rates by 15.3 and 9.2 percentage points, respectively. Second, GPT-6-Astra exhibits promising capabilities in interpreting multi-stage instructions, understanding the environment, and adjusting routes. Third, execution and goal-verification failures persist even with ultra reasoning. Fourth, these results motivate combining general-purpose model capabilities with navigation-specific expertise. Based on these findings, future VLN research should investigate which aspects of instruction interpretation, spatial understanding, and navigation decision-making general-purpose models can handle directly, and where navigation-specific learning can extend their capabilities. This includes exploring how spatial representations, navigation experience, and learned control skills can improve progress tracking, error recovery, and goal verification while preserving the flexibility to adjust routes.

01 / Four findings

Strong performance.
A closer look at how it navigates.

From task-level outcomes to the observations and actions behind them.

01

Strong zero-shot navigation from monocular RGB.

Success reaches 75.7% with medium reasoning and 81.3% with ultra reasoning, without navigation-specific fine-tuning, a trained waypoint predictor, or a pre-built scene map in our evaluation.

Explore the results
02

Promising instruction interpretation, environment understanding, and route adjustment.

Successful trajectories have median nDTW of 90.3–91.0% across ultra runs. Recorded interactions show additional observations, returns to earlier locations, and changes in route or movement direction.

Follow a route adjustment
03

Reliable execution and goal verification remain challenging.

8% of tasks fail in all six evaluations, while 30% have mixed outcomes. Higher reasoning effort does not ensure reliable completion.

Inspect repeated failures
04

Combine general-purpose capabilities with navigation expertise.

Which capabilities can the model provide directly? Where can spatial representations, navigation experience, and learned control skills extend them while preserving flexible route adjustment?

Read the research implications

02 / The evaluation interface

Minimal tools.
The model makes the decisions.

A continuous Codex session connects GPT-6-Astra to the simulator through the Model Context Protocol (MCP).

GPT-6-Astra uses observe() for monocular RGB observations and step(actions) to execute primitive actions through MCP.
Figure 2 The minimal navigation interface.
01 / Look

observe()

Returns the current 512 × 512 RGB view. The model chooses when another observation is useful.

02 / Act

step(actions)

Executes an ordered sequence of basic actions and reports execution counts, remaining budget, and termination status.

Forward 0.25 mTurn left / right 15°Tilt camera up / down 30°STOP

A new RGB view requires another observe() call. Scene maps, reference paths, goal coordinates, global poses, and depth are unavailable through this interface.

100 tasks10 unseen scenes
500 actionsPer-episode budget
2,400 secondsPer-episode time limit
STOP within 3 mRequired for success

03 / Navigation performance

81.3% success.
One monocular camera.

Both reasoning settings use the same R2R-CE-100 tasks, task prompt, and navigation tools.

GPT-6-Astra on R2R-CE-100
Reasoning effortNE ↓ mOSR ↑ %SR ↑ %SPL ↑ %
Medium3.0 ± 0.480.7 ± 2.175.7 ± 1.565.6 ± 2.1
Ultra2.9 ± 0.283.7 ± 1.581.3 ± 2.571.5 ± 1.7

Mean ± s.d. over three runs per reasoning setting. Text and figures use means unless stated otherwise. NE: final goal distance; OSR: entering the goal radius at any point; SR: successful completion; SPL: success weighted by path length.

Explore all reported methods 48 rows

Report Table 1 · reported navigation results
MethodSettingSplitViewNE ↓OSR ↑SR ↑SPL ↑
ScaleVLNTrainedFullPano.4.8—55.051.0
ETPNavTrainedFullPano.4.765.057.049.0
BEVBertTrainedFullPano.4.667.059.050.0
HNRTrainedFullPano.4.467.061.051.0
EnergyTrainedFullPano.4.765.058.050.0
g3D-LFTrainedFullPano.4.568.061.052.0
NavFoMTrainedFullPano.4.672.161.755.3
ABot-N0TrainedFullPano.3.870.866.463.9
OmniNavTrainedFullPano.3.774.669.566.1
Qwen-RobotNav-8BTrainedFullPano.3.578.572.166.6
AstraNav-WorldTrainedFullPano.3.973.967.965.4
NaVidTrainedFullMono.5.749.241.936.5
Uni-NaVidTrainedFullMono.5.653.347.042.7
NaVILATrainedFullMono.5.262.554.049.0
Aux-ThinkTrainedFullMono.5.954.949.741.7
Dynam3DTrainedFullMono.5.362.152.945.7
StreamVLNTrainedFullMono.5.064.256.951.9
DualVLNTrainedFullMono.4.170.764.358.5
InternVLA-N1TrainedFullMono.4.863.358.254.0
D3D-VLPTrainedFullMono.4.767.261.356.1
Image2NavTrainedFullMono.4.072.966.361.5
Qwen-RobotNav-8BTrainedFullMono.4.472.765.759.6
NavGPT-CE-GPT4Zero-shotFullPano.8.426.916.310.2
HSGMZero-shotFullPano.5.458.747.932.8
MapGPT-CE-GPT4oZero-shotR2R-CE-100Pano.8.221.07.05.0
DiscussNav-GPT4Zero-shotR2R-CE-100Pano.7.815.011.010.5
Open-Nav-GPT4Zero-shotR2R-CE-100Pano.6.723.019.016.1
Three-Step Nav-GPT-5Zero-shotR2R-CE-100Pano.5.939.034.029.1
STRIDER-GPT-4oZero-shotR2R-CE-100Pano.6.939.035.030.3
LaViRA-GPT-4oZero-shotR2R-CE-100Pano.6.4 ± 0.2843.3 ± 3.236.0 ± 1.728.3 ± 0.8
LaViRA-Gemini-2.5-proZero-shotR2R-CE-100Pano.6.5 ± 0.2748.7 ± 2.138.3 ± 0.628.3 ± 0.9
O2C-Nav-Gemini-2.5-ProZero-shotR2R-CE-100Pano.5.865.049.330.9
EvoNav-GPT-4oZero-shotR2R-CE-100Pano.6.035.030.024.9
EvoNav-Gemini-2.5-proZero-shotR2R-CE-100Pano.5.051.043.037.8
SmartWay-GPT-4oZero-shotR2R-CE-100Pano.7.051.029.022.5
SmartWay-GPT-5.5Zero-shotR2R-CE-100Pano.5.260.044.035.0
AgenticNav-Gemini-2.5-proZero-shotR2R-CE-100Pano.5.963.049.033.2
AgenticNav-GPT-5.5Zero-shotR2R-CE-100Pano.5.265.055.048.4
SpatialNavZero-shotAuthor-sampledPano.5.266.064.051.1
SpatialAntZero-shotAuthor-sampledPano.4.476.066.054.4
HarnessVLN-GPT-5.5Zero-shot—Pano.4.072.760.843.5
Fast-SmartWay-GPT-4oZero-shotR2R-CE-100F3+P7.7 ± 0.42—27.8 ± 2.2225.0 ± 2.70
CA-NavZero-shotFullMono.7.648.025.310.8
AO-PlannerZero-shotFullMono.7.038.325.516.6
DreamNavZero-shot—Mono.7.141.032.829.0
GC-VLNZero-shotFullMono.7.341.833.616.3
GPT-6-Astra (medium reasoning)Zero-shotR2R-CE-100Mono.3.0 ± 0.480.7 ± 2.175.7 ± 1.565.6 ± 2.1
GPT-6-Astra (ultra reasoning)Zero-shotR2R-CE-100Mono.2.9 ± 0.283.7 ± 1.581.3 ± 2.571.5 ± 1.7

Full: complete validation-unseen split. Author-sampled: correspondence with Open-Nav's subset is unverified. Pano.: panoramic views; Mono.: monocular views; F3+P: three frontal views with initial or on-demand panoramas. LaViRA and Fast-SmartWay report mean ± s.d. over three and four runs. Missing or unverified values are shown as “—”. See Table 1 in the report for publications, score sources, and additional settings.

05 / Failure modes & reliability

More reasoning.
Persistent challenges.

Repeat evaluations distinguish a difficult task from a single unsuccessful execution.

62%

Always succeed

Successful in all six evaluations.

30%

Mixed outcomes

Success varies across evaluations.

8%

Always fail

Unsuccessful in all six evaluations.

100 distinct tasks, each evaluated three times with medium and three times with ultra reasoning.

Three ultra runs have 79, 81, and 84 successes. A task consistency matrix shows 62 always succeeding and 8 always failing. EP176 shows identical views after eight forward actions; EP1133 fails across all six evaluations.
Figure 4 Repeated evaluations distinguish persistent and variable failures. Selected cases show ineffective motion and an incorrect stopping location.
Route following

Plausible landmarks can lead to the wrong place.

In one ultra run, EP513 ends 14.5 m from the goal despite a plausible fireplace and seating match. The same task succeeds in the other two ultra runs.

Execution

Repeated movement does not ensure progress.

In EP176, eight forward actions produce identical RGB views. The episode later exhausts its action budget with 35.4 m navigation error.

Goal verification

Entering the goal area does not ensure a correct stop.

Across the ultra runs, 1–3 failures per run previously enter the success radius. Final location must still be checked against the instruction and route.

06 / Implications for embodied navigation

What can the model provide?
What should navigation research add?

Prior VLN research supplies benchmarks and methods for visual grounding, memory, and action selection. The next question is how this expertise can extend general-purpose models.

Investigate which aspects of instruction interpretation, spatial understanding, and navigation decision-making general-purpose models can handle directly, and where navigation-specific learning can extend their capabilities.

01

Spatial representations

Relate current views to visited places and completed instruction steps. Preserve landmark relations and route history while allowing uncertain judgments to change.

Progress tracking & goal verification
02

Navigation experience

Connect observations and actions with their outcomes through trajectory learning or demonstrations. Test which experiences remain useful across tasks and layouts.

Route choice & error recovery
03

Learned control skills

Provide efficient movement and recovery routines that the model can select or revise. Use execution feedback to detect ineffective motion.

Efficient movement & reliable execution
Preserve the flexibility to observe and adjust routes.

These are directions for future work. Controlled comparisons should test each addition with the same model, tasks, observations, primitive actions, and budgets, while disclosing any added information.

Scope & open questions

What remains to be established

Broader transfer. Repeated runs establish repeatability on this fixed R2R-CE-100 set. Longer instructions, new environments, and other datasets remain to be evaluated.

Practical deployment. High inference cost and latency are current constraints. They may diminish as models and inference systems improve.

Training-data transparency. GPT-6-Astra is proprietary. Whether it was trained on navigation data, and how much, is unknown. Zero-shot here means no navigation-specific fine-tuning in our evaluation.

07 / Citation

Cite this work

The website and downloadable PDF reflect the latest V2 update. The arXiv link provides the public submission record.

BibTeX
@misc{dai2026gpt6astra,
  title = {GPT-6-Astra Lights Up Embodied Navigation:
           Evaluation in Zero-Shot Vision-and-Language
           Navigation in Continuous Environments},
  author = {Guangzhao Dai and Qianru Sun and Qi Wu and Bin Zhu},
  year = {2026},
  eprint = {2609.29861},
  archivePrefix = {arXiv},
  primaryClass = {cs.RO},
  url = {https://arxiv.org/abs/2609.29861}
}
Report figure