REVIEW 8 cited by
PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present PoliFormer (Policy Transformer), an RGB-only indoor navigation agent trained end-to-end with reinforcement learning at scale that generalizes to the real-world without adaptation despite being trained purely in simulation. PoliFormer uses a foundational vision transformer encoder with a causal transformer decoder enabling long-term memory and reasoning. It is trained for hundreds of millions of interactions across diverse environments, leveraging parallelized, multi-machine rollouts for efficient training with high throughput. PoliFormer is a masterful navigator, producing state-of-the-art results across two distinct embodiments, the LoCoBot and Stretch RE-1 robots, and four navigation benchmarks. It breaks through the plateaus of previous work, achieving an unprecedented 85.5% success rate in object goal navigation on the CHORES-S benchmark, a 28.5% absolute improvement. PoliFormer can also be trivially extended to a variety of downstream applications such as object tracking, multi-object navigation, and open-vocabulary navigation with no finetuning.
Forward citations
Cited by 8 Pith papers
-
LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation
LangMap is a human-verified navigation benchmark with 18K tasks spanning scene-, room-, region-, and instance-level goals in real-world 3D scans, covering 414 object categories.
-
ABot-N1: Toward a General Visual Language Navigation Foundation Model
A slow–fast VLN model that routes five navigation tasks through CoT plus image-space pixel goals reaches SOTA on established and new urban benchmarks, including 77.3% POI arrival.
-
MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
A four-camera VLA navigation model trained by distilling multiple RL experts achieves strong simulation performance and qualitative real-world transfer.
-
What Matters in RL-Based Methods for Object-Goal Navigation? An Empirical Study and A Unified Framework
In modular RL-based object-goal navigation, perception quality and test-time strategies dominate performance; policy architecture and observation-space choices contribute little under the tested settings.
-
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation
MTU3D unifies visual grounding and frontier-based exploration in a single transformer, achieving state-of-the-art success rates on HM3D-OVON, GOAT-Bench, SG3D, and A-EQA after large-scale vision-language-exploration p...
-
TrackVLA: Embodied Visual Tracking in the Wild
A single vision-language-action model jointly trained on recognition and tracking data follows described targets at the best reported levels on a public benchmark and transfers zero-shot from simulation to a real quad...
-
Learning Category-level Last-meter Navigation from RGB Demonstrations of a Single-instance
An RGB-only imitation-learning policy trained on one chair generalizes to unseen chairs and environments for last-meter base positioning, though reported success depends on an added heuristic stopping rule and a 0.3 m...
-
Spatially-Enhanced Recurrent Memory for Long-Range Mapless Navigation via End-to-End Reinforcement Learning
A modified recurrent unit with an input-multiplied gate improves spatial memory and long-range mapless navigation success rates by about 23.5% over standard RNNs in simulation and transfers zero-shot to a real robot.
Discussion (0). Continue with ORCID to comment.