A person repairing a machine can be recorded from two very different viewpoints. An egocentric camera can show the hand reaching for the tool, the screwdriver entering the frame, and the precise movements used to tighten a component. An exocentric camera can show where the person is standing, how they move around the machine, which part of the machine they approach and whether another person is assisting them.

Both recordings describe the same repair, but they do not give a model the same evidence. This is the important distinction between egocentric and exocentric data. The viewpoint determines which actions, objects and relationships are visible and at what level of detail.

An egocentric view records an activity from the perspective of the person performing it, typically captured by a device worn on the body. An exocentric view observes the actor from outside the frame using a camera positioned elsewhere in the room or workspace. 

The value in comparing them lies not in which view is objectively better, but in what each one reveals, what it hides and what those differences mean for a model learning from the footage.

What does an egocentric view capture, and where does it fall short?

First-person view of gloved hands using a screwdriver on a mechanical component, highlighting hand movement, point of contact, and the object being acted upon.

An egocentric view gives a model detailed evidence about how a person interacts with the objects around them. Because the recording follows the actor’s own movement, it can capture small, precise actions as they happen. This makes the viewpoint useful for learning how an action unfolds, rather than simply recognising which objects are present.

Fine-grained hand and object interaction

The actor’s hands can be observed as they grasp, release, rotate, reposition or combine objects during a task. A model can therefore associate a specific movement with the object being acted upon, rather than treating the two as unrelated elements in the frame. This distinction matters for any task where having an object in view is not the same as actually using it.

Procedural actions and task sequences

Most tasks involve several actions that have to happen in a certain order. An egocentric recording can show how a person moves from one step to the next and how those steps combine to complete the task. For a model learning from human demonstrations, this reveals which actions come before or after others, and whether the same sequence recurs across similar demonstrations.

Actor-centred context

The camera captures whatever falls within the actor’s field of view, including the immediate workspace and nearby objects involved in the activity. That visibility is also limited by design. Objects or actions outside the frame go uncaptured, and the actor’s own hands or body routinely block parts of the scene.

Where egocentric footage breaks down

The same closeness that makes egocentric footage useful also makes it hard to work with. Constant head and body motion produces motion blur that generic tracking tools struggle with, and an object can leave the frame the moment the wearer turns. Self-occlusion compounds this: the hand performing the action frequently hides the very thing it is acting on.

There is also no stable global frame of reference, since the camera itself never stops moving. A model trained only on egocentric footage can learn what a task looks like from inside it without ever learning where that task takes place, who else is involved or how the actor’s body is positioned while doing it. That gap is exactly what an exocentric view fills.

What does an exocentric view capture, and where does it fall short?

Multi-person interaction, and spatial context around a large mechanical assembly.

An exocentric view observes the actor within the surrounding scene, which lets a model interpret a person’s actions in relation to the people, objects and environment around them. Where an egocentric camera stays locked to the actor’s own perspective, an exocentric camera holds a fixed position and lets the actor move through it.

Whole-body movement and pose

An external camera captures the actor’s entire body, making it possible to observe whether someone is standing, bending, reaching, turning or changing position while performing a task. That broader view connects an action to the posture that accompanies it. For activities where physical positioning affects how a task gets done, this footage reveals patterns that an egocentric recording cannot show.

Spatial and environmental context

An exocentric view shows where an activity takes place and how the actor moves within that space. The footage captures the actor’s position relative to machines, furniture, tools and doorways, along with movement between different areas of the workspace. It can also capture events beyond the actor’s immediate working area, placing individual actions within the wider environment in which they occur.

Multi-person interactions

Activities involving more than one participant introduce information that a single first-person viewpoint cannot capture. An external camera keeps multiple people in view at once, showing where they are, what they are doing and how they respond to one another. During a repair, for example, one person might pass a tool to another while the second person works on the machine, and the footage can show both sides of that exchange.

Where exocentric footage breaks down

Distance costs detail. A camera positioned across a room or mounted above a workspace loses the fine hand movements an egocentric view captures easily, so grasp type, tool orientation and small manipulations are often reduced to a blur or lost behind the actor’s own body. Occlusion works differently here too: rather than the actor’s hand blocking the object, other people, equipment or furniture in the room can block the actor entirely.

A fixed camera also has a fixed field of view. An actor who walks out of frame is simply gone from the recording until they walk back into it, and a task that moves across a large workspace may need multiple synchronised cameras just to stay in view. 

None of this makes exocentric footage less useful, but it does mean the two views trade one kind of blindness for another, rather than one being strictly better than the other.

How does the same task look different from each viewpoint?

The distinction becomes more concrete once the same activity is recorded from both perspectives. The task itself never changes, but each camera captures different information about how it happens.

Cooking

An egocentric recording of someone preparing a meal can show them picking up an ingredient, holding it in place, cutting it and moving it into a pan. An exocentric recording of the same meal can show how the person moves around the kitchen, where the ingredients and utensils are positioned and how the workspace gets used over time. One recording stays close to the knife work, and the other shows the choreography around it.

Tool use and repair

During a repair, an egocentric view can capture which tool gets picked up, how it is held and the precise movements used on a component. An exocentric view can show how the person positions their body around the machine, moves between parts of the workspace and shifts position as the repair progresses. Neither recording is incomplete on its own, but each answers a different question about the same job.

Human-robot interaction

When a person works alongside a robot, an egocentric recording can show what the person sees, where they reach and how they interact directly with the robot. An exocentric recording can show where the person and robot are positioned relative to each other, how they move around one another and what else is happening in the surrounding area. If a person reaches toward a robot to hand it an object, the first-person view captures that reach and the moment of contact, while the external view captures both participants’ movements and positions at once.

Why does paired ego-exo data matter for embodied AI?

A robot learning from human demonstrations needs more than a record of what a person’s hands do. It also needs to know where the person is standing, where relevant objects are located and how the task unfolds across the wider space around them. Paired ego-exo recordings provide exactly that, since the egocentric view supplies the interaction detail while the exocentric view supplies the setting it happens in.

Ego-Exo4D, released by Meta’s FAIR team with 15 university partners, is the clearest example of this at scale. The dataset pairs a wearable egocentric camera with up to four synchronised exocentric cameras across 740 participants in 13 cities, covering 123 real-world scene contexts and 1,286 hours of combined video. Activities range from cooking and bike repair to basketball, bouldering, dance and even COVID rapid antigen testing, all captured as the same task from both viewpoints at once.

What makes a dataset like this useful is not the two camera types on their own, but their synchronisation. When an egocentric frame and an exocentric frame capture the same instant of the same action, a model can learn that two very different-looking recordings describe one event rather than two. That cross-view correspondence is what lets a system trained on this kind of data generalise beyond whichever single camera it happens to be given at inference time.

What does this mean for your annotation pipeline?

Collecting paired ego-exo footage is only the first step. Turning that footage into training data a model can actually use requires annotation layers that neither view demands on its own, and getting these wrong is where most ego-exo pipelines lose accuracy.

Hand-object interaction

Egocentric footage needs frame-level labelling of grasp type, contact points and object state changes, not just bounding boxes around the objects in view. A hand that “holds a screwdriver” and a hand that “turns a screwdriver” look nearly identical in a static frame, but they represent different actions a model needs to distinguish. This annotation layer is where motion blur and self-occlusion do the most damage, so review passes should specifically check frames where the actor’s hand obscures the object.

Temporal segmentation

Both views need consistent action boundaries marked across the full recording, defining exactly where one step ends and the next begins. This matters more in ego-exo pairs than in single-view footage, because the two cameras can make the same action boundary look different: an egocentric camera might show a clean cut between “reaching” and “grasping,” while the exocentric camera captures the same transition as a continuous body movement. Annotators need a shared definition of the boundary before either view gets labelled, not after.

Ego-exo correspondence

This layer doesn’t exist in single-view annotation at all. Every labelled action, object or event in the egocentric stream needs to map to the same moment and the same entity in the exocentric stream, which means annotators are working across two synchronised timelines rather than one. Get this wrong, and the two views stop reinforcing each other; a model trained on misaligned correspondence data can end up learning that egocentric and exocentric footage describe different events rather than the same one.

None of these three layers is optional if the goal is a model that generalises across viewpoints rather than memorising one. Teams that treat ego-exo annotation as “egocentric labelling plus exocentric labelling” tend to end up with two separately usable datasets instead of one that actually pairs. 

Conclusion

Getting your ego-exo data collection and annotation right the first time saves you from re-shooting footage or re-labelling a dataset that doesn’t actually pair. 

Aya Data works with teams collecting and annotating egocentric and exocentric video for embodied AI, robotics and skilled-activity models. To get started, book a 15-minute discovery call to talk through what your dataset needs before you begin recording. 

Frequently Asked Questions

  1. What is the difference between egocentric and exocentric video?

    Egocentric video is captured from the perspective of the person performing an activity, usually through a wearable camera. Exocentric video is captured from outside that perspective, using a camera fixed elsewhere in the room or workspace. The two views describe the same activity but reveal different information about how it happens.

  2. Why not just use egocentric or exocentric data alone?

    Each view has a specific blind spot. Egocentric footage misses whole-body posture, spatial context and anything outside the wearer’s field of view, while exocentric footage loses fine hand-object detail to distance and occlusion. Using both together fills each view’s gap with what the other view captures well.

  3. What is Ego-Exo4D?

    Ego-Exo4D is a large-scale public dataset built by Meta’s FAIR team and 15 university partners, pairing wearable egocentric cameras with synchronised exocentric cameras across 740 participants and 1,286 hours of video. It covers skilled activities like cooking, bike repair, basketball and dance, with both views captured simultaneously for the same task. It remains the most commonly cited reference dataset for ego-exo research.

  4. Does egocentric or exocentric data matter more for robotics?

    Neither view matters more on its own; robotics tends to need both. A robot arm learning to manipulate an object benefits from egocentric-style detail on grasp and contact, while a robot that needs to navigate around people or other equipment benefits from exocentric-style awareness of the wider space.

  5. Do the egocentric and exocentric cameras need to be synchronised during recording?

    Yes. Synchronisation has to happen at capture time, not after the fact, so that the same instant of the same action is timestamped identically across every camera. Attempting to align unsynchronised footage after recording introduces timing errors that undermine the entire point of pairing the views.

  6. What does ego-exo annotation cover that single-view annotation doesn’t?

    Paired footage needs a correspondence layer mapping every labelled action or object in the egocentric stream to the same moment and entity in the exocentric stream. Single-view annotation only ever has to label one timeline; ego-exo annotation has to keep two timelines consistent with each other.

  7. Which industries are adopting paired ego-exo data collection?

    Robotics and embodied AI teams are the most active adopters, since ego-exo pairs are well suited to training models on physical manipulation and navigation together. Sports performance analysis and clinical skills training are also emerging use cases, given the same coaching or assessment value that made skilled-activity domains a focus of datasets like Ego-Exo4D.