
If we make LLMs embodied in a robot, will they have agency?
Sense of Agency in Alter3
Can Alter3 recognize itself when looking in a mirror?



state tokens represent internal state and sensor (e.g. LIDAR) information
action tokens: Autoregressive Control Generation: The final layer of the VLA token pipeline involves action tokens, which are autoregressively generated by the model to represent the next step in motor control. Each token corresponds to a low-level control signal, such as joint angle updates, torque values



consider ’Helix’, a state-of-the-art humanoid robot equipped with a next-generation VLA model. When instructed verbally, “Please take the water bottle from the fridge,” Helix activates its integrated perception system, where a foundation vision-language model (e.g., SigLIP or DINOv2) segments the visual scene to identify the refrigerator, its handle, and the bottle. The language input is processed by an LLM such as LLaMA-4, which tokenizes the instruction and fuses it with the visual context. This fused representation is passed to a hierarchical controller: the high-level policy plans the task sequence (locate handle, pull door, identify bottle, grasp), while a mid-level planner defines motor primitives, such as grasp type and joint trajectories



Explainability important
But LLMs may hallucinate
Here is the complete set of resources, theoretical foundations, code examples, and Google Colab notebook structure formatted in standard Markdown. You can copy and paste this text directly into your course repository or README.md file.
When introducing multimodal embodied AI to students, it is critical to distinguish between Vision-Language Models (VLMs) and Vision-Language-Action (VLA) Models:
| Resource / Framework | Maintainer / Institution | Key Educational Utility | Primary Focus |
|---|---|---|---|
LeRobot (lerobot) |
Hugging Face | Open-source framework providing lightweight dataset loaders, PyTorch policy implementations (ACT, Diffusion Policy, SmolVLA), and physical robot control interfaces. |
| Undergraduate practicals & physical hardware deployment. | ||
| OpenVLA Ecosystem | Stanford Vision & Learning Lab | 7B-parameter generalist manipulation model trained on the Open X-Embodiment dataset. Integrates natively with Hugging Face transformers. |
| Advanced undergraduate/postgraduate demonstration of end-to-end continuous action prediction. |
| Open X-Embodiment | Global Research Consortium | Standardized dataset spanning over 1 million real-world robot trajectories across 22 distinct hardware embodiments. |
| Demonstrating multi-robot cross-embodiment generalization and dataset standardization. |
This self-contained script uses standard PyTorch components to illustrate how visual feature extraction and text embeddings concatenate to output a continuous $7\text{-DoF}$ robot trajectory vector. It runs instantly on any standard CPU.
import torch
import torch.nn as nn
class ConceptualVLA(nn.Module):
"""
Minimal conceptual VLA architecture for classroom demonstration.
Demonstrates how vision tokens and language embeddings fuse
to predict continuous 7-DoF robot motor commands.
"""
def __init__(self):
super().__init__()
# 1. Vision Backbone: Simulates SigLIP / DINOv2 feature extraction
self.vision_encoder = nn.Sequential(
nn.Conv2d(in_channels=3, out_channels=32, kernel_size=5, stride=2),
nn.ReLU(),
nn.AdaptiveAvgPool2d((1, 1)),
nn.Flatten()
)
# 2. Language Backbone: Simulates LLM Token Embedding layer
self.text_embedding = nn.Embedding(num_embeddings=200, embedding_dim=32)
# 3. Action Head: Predicts 7-DoF (dx, dy, dz, droll, dpitch, dyaw, gripper)
self.action_head = nn.Sequential(
nn.Linear(32 + 32, 64),
nn.ReLU(),
nn.Linear(64, 7)
)
def forward(self, image_tensor, text_tokens):
# Extract features from both input modalities
visual_feats = self.vision_encoder(image_tensor)
text_feats = self.text_embedding(text_tokens).mean(dim=1)
# Multimodal fusion via feature concatenation
fused_context = torch.cat([visual_feats, text_feats], dim=-1)
# Output continuous 7-DoF motor control vector
action_vector = self.action_head(fused_context)
return action_vector
if __name__ == "__main__":
# Instantiate toy model
vla = ConceptualVLA()
# 1. Simulate RGB Camera Frame (Batch=1, Channels=3, Height=224, Width=224)
simulated_camera_frame = torch.randn(1, 3, 224, 224)
# 2. Simulate Tokenized Instruction: "Pick up the red block"
simulated_prompt_tokens = torch.tensor([[12, 45, 89, 102]])
# 3. Forward pass
predicted_action = vla(simulated_camera_frame, simulated_prompt_tokens)
labels = ["dx (m)", "dy (m)", "dz (m)", "droll (rad)", "dpitch (rad)", "dyaw (rad)", "gripper (0-1)"]
values = predicted_action.detach().numpy()[0]
print("=" * 50)
print(" SIMULATED VLA PREDICTED 7-DoF ACTION OUTPUT")
print("=" * 50)
for label, val in zip(labels, values):
print(f" {label:<15}: {val:+.4f}")
print("=" * 50)
bitsandbytes)Standard OpenVLA in 16-bit precision requires over $16\text{ GB}$ of VRAM. Applying 4-bit quantization via bitsandbytes reduces the VRAM requirement down to approximately $5\text{ GB}$ to $7\text{ GB}$, making real inference accessible on consumer GPUs or Google Colab T4 runtimes.
import torch
from PIL import Image
import requests
from transformers import AutoModelForVision2Seq, AutoProcessor, BitsAndBytesConfig
# 1. Configure 4-bit quantization to fit within limited GPU VRAM
quantization_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_compute_dtype=torch.bfloat16
)
print("Loading OpenVLA-7B model and processor...")
# 2. Load OpenVLA Processor and Quantized Weights
processor = AutoProcessor.from_pretrained("openvla/openvla-7b", trust_remote_code=True)
vla_model = AutoModelForVision2Seq.from_pretrained(
"openvla/openvla-7b",
quantization_config=quantization_config,
low_cpu_mem_usage=True,
trust_remote_code=True
)
# 3. Download sample tabletop camera frame
img_url = "https://raw.githubusercontent.com/openvla/openvla/main/assets/bridge_example.png"
image = Image.open(requests.get(img_url, stream=True).raw).convert("RGB")
# 4. Formulate task instruction
prompt = "In: What action should the robot take to pick up the cucumber? \nOut:"
# 5. Predict action vector
inputs = processor(prompt, image).to("cuda")
predicted_action = vla_model.predict_action(**inputs, unnorm_key="bridge_orig", do_sample=False)
print("\n" + "=" * 50)
print(" OPENVLA REAL-WORLD PREDICTED ACTION VECTOR")
print("=" * 50)
print("Action Vector [x, y, z, roll, pitch, yaw, gripper]:")
print(predicted_action)
print("=" * 50)
.ipynb Ready)Below is the cell-by-cell Markdown structure for setting up a complete Google Colab practical session for students.
Module: Autonomous Mobile Robotics & VLA Architectures
Runtime Requirements:
Runtime -> Change runtime type -> T4 GPU)!pip install -q torch torchvision transformers accelerate bitsandbytes pillow timm
This section demonstrates the internal components of a VLA:
import torch
import torch.nn as nn
class MinimalVLA(nn.Module):
def __init__(self):
super().__init__()
self.vision_backbone = nn.Sequential(
nn.Conv2d(3, 32, kernel_size=5, stride=2),
nn.ReLU(),
nn.AdaptiveAvgPool2d((1, 1)),
nn.Flatten()
)
self.text_embedding = nn.Embedding(200, 32)
self.action_head = nn.Sequential(
nn.Linear(32 + 32, 64),
nn.ReLU(),
nn.Linear(64, 7)
)
def forward(self, image_tensor, text_tokens):
visual_feats = self.vision_backbone(image_tensor)
text_feats = self.text_embedding(text_tokens).mean(dim=1)
fused_context = torch.cat([visual_feats, text_feats], dim=-1)
return self.action_head(fused_context)
# Execution
model = MinimalVLA()
simulated_image = torch.randn(1, 3, 224, 224)
simulated_tokens = torch.tensor([[12, 45, 89, 102]])
predicted_actions = model(simulated_image, simulated_tokens)
labels = ["dx (m)", "dy (m)", "dz (m)", "droll (rad)", "dpitch (rad)", "dyaw (rad)", "gripper (0-1)"]
print("=" * 50)
print(" CONCEPTUAL VLA 7-DoF PREDICTION OUTPUT")
print("=" * 50)
for label, val in zip(labels, predicted_actions.detach().numpy()[0]):
print(f" {label:<15}: {val:+.4f}")
print("=" * 50)
This section loads OpenVLA-7B using 4-bit quantization (bitsandbytes). 4-bit quantization reduces GPU memory usage from $16\text{ GB}$ down to $\sim 5\text{ GB}$, fitting within the free Google Colab T4 GPU quota.
import torch
from PIL import Image
import requests
from transformers import AutoModelForVision2Seq, AutoProcessor, BitsAndBytesConfig
# 1. Configure 4-Bit Quantization
quantization_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_compute_dtype=torch.bfloat16
)
print("Loading OpenVLA-7B processor and quantized model...")
# 2. Load Model & Processor
processor = AutoProcessor.from_pretrained("openvla/openvla-7b", trust_remote_code=True)
vla_model = AutoModelForVision2Seq.from_pretrained(
"openvla/openvla-7b",
quantization_config=quantization_config,
low_cpu_mem_usage=True,
trust_remote_code=True
)
# 3. Download Test Image
img_url = "[https://raw.githubusercontent.com/openvla/openvla/main/assets/bridge_example.png](https://raw.githubusercontent.com/openvla/openvla/main/assets/bridge_example.png)"
image = Image.open(requests.get(img_url, stream=True).raw).convert("RGB")
# 4. Define Prompt
prompt = "In: What action should the robot take to pick up the cucumber? \nOut:"
# 5. Predict Action Vector
inputs = processor(prompt, image).to("cuda")
predicted_action = vla_model.predict_action(**inputs, unnorm_key="bridge_orig", do_sample=False)
print("\n" + "=" * 50)
print(" OPENVLA PREDICTED 7-DoF ACTION VECTOR")
print("=" * 50)
print(predicted_action)
print("=" * 50)