01
Vision-Language-Action models
Policies that take in camera images and a written instruction and output motor commands directly. I want to understand how well these models ground language in what a robot actually sees, and how far they transfer to a new robot with little data.
What I want to explore
- Grounding instructions in visual scenes
- Fine-tuning on small, real robot datasets
- Generalising across different robot bodies



