Agents Papers

H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning

Zhang et al.

A three-level hierarchy of predictive world models lifts Visual AntMaze success from 18% to 73% with less compute.

H-JEPA stacks joint-embedding predictive world models into a hierarchy and plans from the top level down, with higher levels setting goals for lower ones.

A three-level hierarchy raises success on the Visual AntMaze benchmark from 18 to 73 percent while using less compute. It was the most upvoted paper of the week on alphaXiv.