intrin. Sydney, Australia

Computer science · Embodied intelligence

Haoran Liu.

刘浩然

From perception
to action.

I’m a computer science master’s student at the University of Sydney, exploring how machines see the world, build models of it, and learn to act.

OBSERVE / IMAGINE / ACT
From observations to a model of the world A conceptual illustration of three layered spatial grids, connected by a path through observation, imagination and action. 01 / SEE02 / MODEL03 / ACT
A curiosity that connects pixels to the physical world.
Areas of interest

Computer vision / World models / Robot learning / Multimodal AI

01 / Selected work

Ideas, made tangible.

A few things I’ve built
and questions I’m exploring.

VELOREN / RECORDED GAMEPLAY02
Recorded agent gameplay · VelorenFrozen JEPA agent · Continuous capture at 30 fps.
Selected demo, not overall task success. No audio.

Research in progress

LLCL

What can an agent learn to anticipate?

Exploring world models for multi-step decisions in games, with generalization to new layouts as a central research question.

  • DreamerV3
  • ViZDoom
  • World models

Code not public

Inside the research

The current work investigates a DreamerV3 world model with an actor–critic policy in ViZDoom, following earlier exploration of JEPA and Veloren.

The video shows a selected run of the frozen JEPA controller in Veloren. Evaluation is ongoing; performance on training maps does not establish generalization to unseen layouts.

MUJOCO / DATASET CAPTURE03
Real simulation footage · 3× playbackTwo synchronized views of a robot drawing on a whiteboard. Dataset capture, not a policy benchmark. No audio.

Simulation practice

Robot learning, in simulation

A space to connect perception and action.

Hands-on MuJoCo simulation work in the context of vision–language–action datasets, alongside experience with ROS 2 and RoboMaster.

  • MuJoCo
  • VLA datasets
  • ROS 2

Code not public

About this exploration

This work reflects practical simulation experience and an interest in how embodied datasets connect observations, instructions and actions.

It is an exploration of the tooling and problem space; no policy-training or benchmark results are claimed.

02 / A little about me

Curiosity, with
a hands-on habit.

I’m drawn to the space between understanding a scene and doing something useful in it.

At the University of Sydney, I’m pursuing a master’s in computer science. My interests connect computer vision, world models, robot learning and multimodal AI — from the geometry of aerial images to agents that learn through interaction.

I like working through the full path from an idea to a working system: the models, the implementation details, and the experiments that reveal what still needs work.

Also on my workbench

Speech & video understanding. A subtitle pipeline combining ASR, speaker-aware timing and translation, alongside exploration of livestream understanding and clips. Subtitle pipeline on GitHub

Human–computer interaction. An exploration of bee pollination through sensing, haptic feedback and mechanical flower prototypes.