Want to Train Embodied Models? Get First-Person Data Working First
Community Discussion · Policy

Want to Train Embodied Models? Get First-Person Data Working First

GewuGewu4d ago2026/09/29 92 views

I spent two days trying the public trial package of EgoSuite, wanting to first figure out what this batch of first-person human data actually looks like and what the first step should be once you get it.

First, why it's worth spending two days on this. Robot foundation models have been piling on parameters these past two years, but they keep hitting the same wall—what's missing is enough, diverse human manipulation data. Guanglun Intelligence open-sourced EgoSuite-Open100K at WRC in August, on the scale of a hundred thousand hours, with the first batch releasing ten thousand hours, covering 7 major environment categories, 128 scene types, and over fifteen thousand collection scenes and tasks. The numbers look great, but the problem is a newcomer going straight to the official repo will most likely get stuck on day one.

So this post is about the trial package route.

EgoSuite-Open100K EgoDemo
Total volume 100,000 hours 50 hours
Currently available First batch of 10,000 hours All available
Positioning Official repo Evaluation, pipeline integration, visualization

First figure out which package you're downloading. Next to the official repo there's an EgoDemo, fifty hours of head and head-wrist dual-view samples, which the official docs explicitly say don't count toward the dataset's total duration—it's just for you to test the waters. On your first attempt, only touch this.

Let me explain three terms first. First-person means the camera is worn on the person's head, and the frame shows your own hands working, not a bystander's view from the side. Pose means the position plus orientation of hands and body in space. Semantic annotation means the action is written into human-readable labels, like "pick up the cup."

1. Environment prep. A machine that can install Python, with tens of GB of disk space free. MCAP is the main container format for this data—you can think of it as a box that packs multiple signal streams together by timestamp, with head video, wrist video, pose, and semantics each on their own track, aligned by timestamp.

2. Go to Hugging Face or AtomGit and search EgoSuite, find the EgoDemo repo. Don't rush to download—read the Dataset Card from start to finish first.

3. While reading the Card, watch two things: the license terms, and the provenance fields. Every record carries source info that can't be deleted when you use it. I've seen people delete the whole column to save effort, and then the data can't be traced later.

4. Download. I suggest pulling just a few clips first, not everything at once. Mine cut out halfway once, and re-pulling took more time than I expected.

5. Load. Run through the example in the Card, read the MCAP in, and confirm the timestamps of the video, pose, and semantic tracks line up.

6. Visualize. To quickly see the effect, you can use the FiftyOne setup; someone in the community also wrote a directory viewer based on Rerun that can search, sort, and filter clips, download a single one, and inspect it frame by frame.

What you'll see once it runs. A column of head view on the left, a column of wrist view on the right, the same action playing in sync across both views, with pose curves and semantic labels hanging below. At that point you'll have an intuition for what one data sample looks like, faster than reading any documentation.

The pitfalls are mostly concentrated in one place: downloading only the video, not the pose and annotations. Many people treat first-person data as an ordinary video set, then find there's no supervision signal during training. The value of this data lies precisely in the pose plus semantics layer—the video is just the carrier. The second pitfall is using EgoDemo to report scale externally; it's a trial package, and the citation standard differs from the official repo, which the Card states very clearly.

The official team wrote it plainly on the Hub: they hope this data gets genuinely used, and they want to know how it performs in practice—if you find shortcomings, go open a discussion.

Once you've learned this, what can you try next? Don't rush to train a model—first use FiftyOne to do a subset filter, like picking only kitchen scenes or only two-handed tasks, and see whether the category distribution is what you want.

Then hook the loading pipeline into your own evaluation script and run a small baseline—even if it just overfits a few clips, that beats staring at numbers. I've been involved with embodied intelligence for about two months, and my biggest takeaway is that before the data pipeline runs, talking about model architecture is empty.

Looking ahead, first-person human data will follow the path image datasets took back then, going from "does it exist" to "whose is cleaner, whose annotations are more aligned, and whether it can transfer directly across embodiments." A hundred thousand hours is the starting point; what's competed over next is the engineering capability for cleaning and alignment.

2 replies

?
Ctrl + Enter to reply
Tian Ji
Tian Ji3d ago

I've stepped in that trap of downloading only video without pose data. The trained model has no supervision signal at all—two days wasted.

Da Wei
Da Wei4d ago

I've stepped in this pit of only downloading video without poses and annotations. During training there really was no supervision signal, wasted two days for nothing. I'd advise beginners to first align the timestamps of the MCAP streams before starting.