GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

GR00T N1은 VLM(System 2)이 '무엇을 할지' 이해하고, 별도의 DiT action expert(System 1)가 cross-attention으로 그 정보를 읽어 flow matching으로 action을 만드는 dual-system VLA다. 구조, data pyramid, 학습·추론, 결과, 그리고 DreamZero·Dyna와의 위치 차이를 정리한다.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference


GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

이 글의 그림은 위 논문에서 가져왔다.

Overview


논문 Figure 3: GR00T N1 model architecture

  • 핵심 구조는 VLM(System 2) + DiT action expert(System 1).
  • VLM은 image + language를 해석해 vision-language token $\phi_t$를 만든다.
  • DiT는 robot state $q_t$와 noisy action chunk $A_t^\tau$를 self-attention으로 처리하고, 중간중간 cross-attention으로 VLM token을 읽는다.
  • Cross-attention에서 $Q$는 DiT의 state/action hidden, $K,V$는 VLM의 vision-language token이다.
  • 즉 action branch가 “지금 어떤 semantic/visual 정보가 필요한가?”를 Query하고 VLM representation에서 필요한 정보를 가져온다.
  • 속도가 다르다: VLM(System 2)은 10Hz, DiT(System 1)는 120Hz로 action을 낸다.
  • 모델 크기는 전체 2.2B(그중 VLM 1.34B). action 16개(H=16) 한 chunk를 뽑는 데 L40 GPU에서 63.9ms 걸린다.

DATA


  • GR00T N1의 pretraining data는 3층 data pyramid로 보면 된다.
    1. Human video / web-scale data: 양은 많지만 robot action label이 없음.
    2. Synthetic data: simulation trajectory + neural trajectory.
    3. Real robot data: 가장 비싸지만 실제 control signal이 존재.
  • Real robot data에는 Fourier GR-1 중심 teleoperation data와 Open X-Embodiment 등 여러 embodiment가 포함된다.
  • Human video에는 ground-truth action이 없으므로 Latent Action Pretraining(LAPA)로 pseudo latent action을 만든다.
  • Neural trajectory는 image-to-video model로 trajectory를 생성한 뒤 inverse-dynamics model 또는 latent-action model로 pseudo action을 붙인다. 이렇게 직접 모은 teleoperation data 88시간을 약 827시간(약 10배)으로 늘렸다.
  • 목적은 다양한 embodiment와 data source를 하나의 shared robot foundation policy pretraining에 묶는 것이다.

논문 Figure 1: data pyramid

Training method


Architecture / Attention


  1. VLM(System 2): image + language → vision-language tokens $\phi_t$. NVIDIA Eagle-2 기반(SigLIP-2 이미지 인코더 + SmolLM2 LLM). 이미지 한 장은 224×224 해상도로 64 token이 된다.
    • 중요: VLM의 마지막 layer가 아니라 중간 layer(12번째 LLM layer) 출력을 쓴다. 논문 실험에서 이쪽이 추론도 빠르고 성공률도 높았다. (DiT4DiT가 video DiT의 마지막이 아닌 layer 18을 쓴 것과 같은 이야기)
  2. State / Action projection: embodiment마다 state/action dimension이 다르므로 embodiment-specific MLP로 shared embedding dimension에 projection.
  3. DiT(System 1): cross-attention block과 self-attention block이 번갈아 쌓인 구조(Flamingo 방식). diffusion timestep은 AdaLN으로 넣는다.
    • self-attention: $[q_t,A_t^\tau]$ 내부에서 state와 noisy action이 서로 정보를 주고받는다.
    • cross-attention: DiT hidden이 Query, VLM token $\phi_t$가 Key/Value.
\[Q=H_{DiT}W_Q,\qquad K=\phi_tW_K,\qquad V=\phi_tW_V\] \[\mathrm{CrossAttn}=\mathrm{softmax}\left(\frac{QK^T}{\sqrt d}\right)V\]
  1. Action Decoder: 마지막 action hidden을 embodiment-specific MLP decoder로 continuous action으로 변환.

Loss function


  • GR00T N1의 action generation은 flow matching이다.
  • Ground-truth action chunk:
\[A_t=[a_t,a_{t+1},\ldots,a_{t+H-1}]\]
  • noise $\epsilon$와 interpolation time $\tau$를 사용해 noisy action을 만든다.
\[A_t^{\tau}=\tau A_t+(1-\tau)\epsilon\]
  • DiT는 noisy action을 clean action 쪽으로 이동시키는 velocity field를 예측한다.
\[V_\theta(\phi_t,A_t^{\tau},q_t)\approx \epsilon-A_t\]
  • → 부호 주의: 위 보간식을 τ로 미분하면 정답 방향은 A_t − ε (noise → action)인데, 논문은 π₀ 논문(v1–v3)과 같은 표기로 ε − A_t라고 적었다. 추론 업데이트 A^(τ+1/K) = A^τ + (1/K)·V 도 V가 A_t − ε 방향이어야 noise에서 action으로 간다. 방향 부호만 다를 뿐 같은 얘기다 (자세한 건 π₀ 글 참고).
  • 학습 때 τ는 π₀처럼 Beta 분포로 샘플링한다.
  • inference에서는 random noisy action chunk에서 시작해 이 velocity를 반복 적용하여 continuous action chunk를 복원한다. 반복 횟수는 K=4 step이면 충분했다고 한다.

Pre-training


  • 다양한 embodiment + real robot + synthetic + human video를 한 모델에 섞어 대규모 pretraining한다.
  • pre-training 때도 VLM의 language 쪽(text tokenizer/LLM)은 freeze하고, vision encoder와 DiT만 학습한다.
  • robot trajectory는 실제 action과 latent action을 flow-matching target으로 사용할 수 있다.
  • human video에는 robot action이 없으므로 LAPA가 만든 latent action을 target으로 사용한다.
  • neural trajectory에는 learned latent action 또는 inverse-dynamics model이 예측한 pseudo action을 붙인다.
  • 목표는 특정 robot 하나의 task memorization보다 여러 embodiment와 환경에 대한 general robot prior를 만드는 것이다.

Post-training


  • pretrained GR00T N1을 특정 embodiment / downstream task dataset으로 fine-tuning한다.
  • pre-training과 마찬가지로 language component는 frozen 상태로 두고 나머지 vision/action 쪽을 fine-tune한다.
  • task-specific real data가 적을 때는 neural trajectory를 추가 생성해 real trajectory와 co-training할 수 있다.
  • 핵심 목적은 대규모 pretraining의 generalization을 유지하면서 적은 downstream data로 target robot/task에 빠르게 적응시키는 것이다.

Inference


  1. 현재 camera image + language instruction → VLM → $\phi_t$.
  2. robot state $q_t$와 random noisy action chunk $A_t^0$를 DiT에 입력.
  3. DiT self-attention으로 state/action 관계를 처리하고, cross-attention으로 VLM token을 참고.
  4. flow-matching velocity를 반복 적용해 noise → continuous action chunk로 denoise.
  5. Action Decoder가 target embodiment의 실제 command dimension으로 변환.

핵심 흐름: Vision + Language → VLM semantic representation → Cross-Attention → DiT action denoising → continuous action chunk.

결과


  • 시뮬레이션 (RoboCasa, DexMimicGen, GR-1 평균 성공률): BC-Transformer 26.4%, Diffusion Policy 33.4%, GR00T N1 45.0%.
  • 실제 GR-1 휴머노이드 (pick-and-place, 관절 물체, 산업 작업, 양손 협응 평균):
    • Diffusion Policy: 데이터 10% 10.2% / 전체 데이터 46.4%
    • GR00T N1: 데이터 10% 42.6% / 전체 데이터 76.8%
  • 즉 GR00T N1은 데이터 10%만으로도 Diffusion Policy가 전체 데이터로 낸 성적에 가깝다. pretraining의 data efficiency가 보이는 부분이다.

기존 VLA와 무엇이 다른가?


  • GR00T N1도 본질적으로는 VLA다. WAM처럼 future-video latent를 직접 생성하는 구조는 아니다.
  • 가장 큰 차별점은 dual-system architecture다.
    • System 2(VLM): semantic reasoning / perception 담당.
    • System 1(DiT): high-frequency continuous control 담당.
  • 많은 VLA가 image/language/action token을 하나의 multimodal Transformer 안에서 직접 섞는 방식이라면, GR00T N1은 VLM representation과 action-generation backbone을 구조적으로 분리하고 cross-attention으로 연결한다.
  • action은 autoregressive discrete token이 아니라 flow-matching continuous action chunk로 생성한다. (단 이건 RT-2·OpenVLA 같은 token 방식과 비교할 때의 차이다. π₀도 flow matching + action expert라서, π₀와 비교하면 진짜 차이는 VLM과 action 모듈을 cross-attention으로만 연결하고(π₀는 같은 attention 안에서 섞는다), VLM(10Hz)과 DiT(120Hz)가 다른 속도로 따로 돈다는 점이다.)
  • real + synthetic + human video를 latent action / neural trajectory를 이용해 하나의 pretraining corpus로 묶는 것도 중요한 특징이다.

DreamZero / Dyna와의 위치 차이


  • GR00T N1: 현재 image/language semantic representation → action. 미래 video latent를 생성하거나 action에 직접 넣지 않음.
  • Dyna-2 main co-training: video prediction objective로 representation을 dynamics-aware하게 만들지만 future-video latent를 action에 직접 넣지 않음.
  • DreamZero: future-video latent + future-action latent를 같은 Joint DiT에서 같이 denoise.
  • 따라서 GR00T N1은 우리가 정리한 분류에서 action-centric VLA 쪽에 가깝고, explicit future imagination을 action generation에 넣는 WAM은 아니다.

장점 / 한계


장점

  • 모듈 분리가 명확: VLM의 semantic reasoning과 DiT의 continuous control 역할이 분리되어 구조 해석과 adaptation이 쉽다.
  • Cross-attention conditioning: action branch가 VLM representation에서 필요한 semantic/visual 정보만 읽어올 수 있다.
  • Continuous control: flow matching + action chunking이 부드러운 연속 motor command 생성에 적합하다.
  • Cross-embodiment: embodiment-specific MLP를 두면서 핵심 backbone을 공유해 다양한 robot morphology를 한 foundation model에 학습할 수 있다.
  • Data scalability: real robot뿐 아니라 synthetic trajectory와 human video까지 latent action으로 활용한다.

한계

  • Explicit world prediction이 없음: DreamZero/DiT4DiT처럼 future scene dynamics를 latent로 직접 상상해 action에 넣는 구조는 아니다.
  • VLM representation 의존: perception / grounding이 틀리면 action DiT도 잘못된 conditioning을 받는다.
  • Iterative denoising cost: 한 번의 deterministic regression보다 action generation inference compute가 더 든다.
  • Post-training specialization: 특정 task data에 강하게 fine-tune하면 pretraining에서 얻은 일부 general behavior가 약해질 수 있다.

한 줄 정리: GR00T N1 = VLM이 ‘무엇을 해야 하는지’ 이해하고, 별도의 DiT action expert가 cross-attention으로 그 정보를 가져와 ‘어떻게 움직일지’를 flow matching으로 생성하는 dual-system VLA.