Egocentric data refers to information captured from a person’s point of view, typically through a wearable or head-mounted camera. A recording of an individual preparing a meal captures their hands, the ingredients used and the tools engaged. The same recording may also capture a family member entering the room, a phone screen left visible on the counter or an address printed on a delivery package.
This is the central operational challenge in egocentric data collection. Unlike a fixed camera recording a defined space, a wearable device travels with the participant throughout the task, and the resulting footage captures everything within its field of view rather than only the activity being studied. A protocol that specifies what should be recorded, without equal attention to consent and privacy, leaves this additional exposure unmanaged by default rather than addressed by design.
The stakes attached to this gap have grown alongside demand for physical AI training data. Recent reporting on egocentric collection practices in industrial and domestic settings has highlighted instances where participants were recorded without a clear understanding of how their footage would be used or how far their consent was intended to extend once collection was complete. Establishing a defined protocol at the outset is therefore not a procedural formality. It is the mechanism that determines whether the resulting dataset can be used responsibly.
Why Consent Cannot Be a One-Time Form

A signed consent form is often treated as the end of the compliance process rather than its beginning. In practice, the footage a participant consents to collect will pass through several further stages before it reaches its final use: review, redaction, annotation and, ultimately, model training. Each of those stages is a point at which the original consent needs to remain valid and applicable.
Consent materials should therefore specify considerably more than permission to record. They should state explicitly whether the footage may be used for annotation, whether it may be reviewed by third-party specialists, how long it will be retained, whether it may be included in datasets shared with other organisations and what a participant’s withdrawal actually affects once granted. Where any of these terms is left undefined, the organisation collecting the data is left to interpret its own permissions after the fact, which is not a position any compliance-conscious buyer should accept from a vendor.
This distinction matters most in regulated or reputationally sensitive sectors, where the same discipline already applies to medical data annotation under HIPAA and GDPR. Egocentric data collection, though a newer field, carries comparable exposure. A robotics or embodied AI programme that treats consent informally is accumulating the same category of risk that healthcare AI programmes have already been required to formalise.
Documentation should record which version of the consent materials a participant accepted, since protocols are frequently revised as a collection programme matures. Where a later version introduces a new use case for the footage, participants who accepted an earlier version have not agreed to it, and the dataset should reflect that distinction rather than assume blanket coverage.
What Happens When Someone Who Never Consented Is in the Frame
A participant’s consent covers their own presence in a recording. It does not cover anyone else who ends up in the frame, and with egocentric footage, that happens often. A colleague walking past, a family member entering the room or a stranger passed on the street can all appear in footage the participant never meant to capture beyond their own task.
There are two ways to handle this once it’s spotted. The clip can be redacted, with the bystander’s face or other identifying details removed while the rest of the footage stays in the dataset. Or the clip can be excluded entirely. This choice should not be left to whoever happens to be reviewing the footage that day. It needs a clear policy: how much of the bystander is visible, whether redaction actually hides their identity, and whether the setting itself, a specific home or workplace, still gives them away even with a blurred face.
The decision also needs to be recorded, not just made. For every clip that is redacted or excluded, the dataset should show which choice was made and under which version of the policy. If a client or regulator later asks how a specific piece of footage was handled, this record is what proves the process was followed consistently rather than decided on the spot.
This review does not have to sit with the team that collected the footage. When a specialist human-in-the-loop team handles redaction and eligibility review against a defined policy, the result is a documented, auditable step between raw recording and anything an annotator actually sees.
Building the Collection Protocol – Egocentric Data Collection
A collection protocol needs three things settled before recording begins: what to capture, how to capture it and proof that both instructions actually work in practice.
The starting point is defining the data itself. If the task is meal preparation, the protocol should state exactly what needs to be captured – hand movements, tool use, interaction with the workspace – rather than leaving participants to interpret the brief themselves. A narrower scope also reduces the amount of unrelated footage the team has to review and redact later, since anything outside the defined task is more likely to capture information the project never needed.
From there, the protocol sets the recording guidelines: expected duration, location, equipment and, importantly, when recording should stop. Participants need clear instructions for pausing or ending a recording, such as entering a private space or encountering someone who does not want to appear on camera. Leaving this judgement to the participant in the moment, without prior guidance, is how avoidable footage ends up in the dataset.
The final step is a pilot collection, run with a small group before the full programme begins. This tests whether the equipment performs as expected and, more importantly, whether participants actually understand and follow the instructions as written. A protocol that reads clearly on paper often reveals gaps once a real participant is wearing the camera. Identifying those gaps in a pilot of a handful of participants is inexpensive. Identifying them after several hundred hours of footage have already been collected is not.
Aya Data structures new egocentric and physical AI data acquisition programmes around this same sequence, and offers a free pilot to validate a protocol at a small scale before a client commits to full collection.
What Happens When Consent Is Withdrawn After the Fact
Giving participants the right to withdraw consent is only useful if the organisation holding the data knows what withdrawal actually requires. The answer depends on how far the footage has already travelled.
If the footage is still raw and unreviewed, withdrawal is straightforward. It gets deleted, and the record shows it was removed at the participant’s request.
If the footage has already been annotated, withdrawal means pulling the annotated data from the dataset before it is used in a training run. This is only possible if the dataset was built with traceability from the start, meaning every annotation can be traced back to the participant and the recording it came from. Without that traceability, a withdrawal request cannot be honoured properly, since nobody can isolate which parts of the dataset need to come out.
The hardest case is footage that has already been used to train a model. At that point, the specific data cannot simply be deleted from the model itself. What withdrawal should trigger here is a documented decision: whether the model needs to be retrained without that participant’s data, whether the exposure is assessed and accepted with a clear rationale, or whether some other mitigation applies. The point is that this decision gets made deliberately and recorded, not skipped because it is inconvenient.
This is why consent and data governance need to be designed together from the outset, rather than treated as separate concerns. A collection protocol that cannot answer what withdrawal means at each stage of the pipeline is not actually offering participants a meaningful right, only a symbolic one.
Access Tiers – Who Should See Raw Footage vs. Reviewed Data
Not everyone working with an egocentric dataset needs to see the same version, and treating access as a single yes-or-no permission understates the risk involved.
Raw footage, unreviewed and unredacted, should sit in the smallest access tier. Only the team responsible for privacy review and redaction should be able to view it, since this footage is most likely to contain bystanders, private conversations or incidental personal information that was never meant to be part of the dataset.
Once footage has passed privacy review, it can move into a broader working tier. Annotators need access to this version to carry out their work, but they should never need access to the raw, unreviewed material behind it. Structuring access this way means a labelling error or a compromised account exposes only reviewed footage, not the full raw archive.
A further tier applies to any footage or derived samples used for demonstration, client reporting or external sharing. This should be the most restricted version of all, limited to material that has been reviewed, redacted and specifically cleared for that purpose.
Defining these tiers in advance, rather than granting broad access by default and narrowing it only when a problem arises, is what allows an organisation to demonstrate, not just claim, that privacy protections were built into the workflow. Aya Data applies this tiered structure across its human-in-the-loop review process, so raw recordings remain restricted to the reviewers who clear them before any footage reaches an annotator.
Preparing Egocentric Data for Annotation

Once recording is complete and footage has passed privacy review, it still needs preparation before it reaches an annotator. This stage begins with identifying anything unsuitable for use: incomplete recordings, duplicates or footage that failed the earlier review and should not have entered this stage at all.
The usable recordings are then organised to the project’s requirements. This typically means consistent file naming, grouping by activity or task type and attaching the metadata annotators need to work efficiently, including which consent version applies to that footage and which access tier it belongs to.
This is also the stage where the earlier governance decisions become visible in practice rather than remaining policy on paper. A dataset arriving at annotation with clear labelling, documented consent status and a defined access tier is one an annotation team can work through efficiently and defend later if a client or regulator asks how a specific piece of footage was handled. A dataset arriving without that structure pushes the same questions downstream, when they are considerably harder to answer.
Conclusion
The value of egocentric data lies in how closely it reflects real-world activity. That same quality is what makes it difficult to control once recording begins, since a camera worn through a task inevitably captures more than the task itself.
Good egocentric data annotation and collection is not defined by how much footage is gathered. It is defined by whether the organisation collecting it can answer, at every stage, who consented, what they consented to, who else appears in the footage without having agreed to it, and what happens if that consent is later withdrawn. A protocol that cannot answer these questions with documentation, not just policy, has not actually solved the problem it set out to address.
At Aya Data, we build this governance into the collection and annotation workflow from the outset, not as a review step applied after the fact. Every stage, from pilot collection through privacy review to tiered access control, is designed to produce a dataset an organisation can stand behind if a client, partner or regulator asks how it was handled. If your team is building an egocentric or physical AI data collection programme and wants that governance in place from the start, book a 15-minute discovery call to review your current protocol against it.
Frequently Asked Questions
What is egocentric data collection?
Egocentric data collection captures footage from a person’s point of view, typically using a wearable or head-mounted camera, recording their hands, tools and surroundings while they carry out a task.
Why is consent harder to manage in egocentric data collection than in other types of recording?
A wearable camera travels with the participant and captures everything in its field of view, not only the task being recorded. This means consent must account for bystanders, private information, and later use of the footage, not just the recording itself.
What happens if someone who never consented appears in the footage?
The footage is either redacted, with identifying details removed, or excluded from the dataset entirely. This decision should follow a documented policy rather than individual reviewer discretion, and the outcome should be recorded for each affected clip.
Can a participant withdraw consent after their footage has already been annotated?
Yes, but honouring that request depends on traceability. If every annotation can be traced back to the participant and the recording it came from, the affected data can be identified and removed before it reaches a training run.
Why does egocentric data need tiered access control?
Raw, unreviewed footage carries the highest privacy risk, since it may still contain bystanders or private information. Restricting raw footage to the reviewers responsible for clearing it, and giving annotators access only to reviewed material, limits exposure if an account is compromised or an error occurs.
Why run a pilot before starting full-scale egocentric data collection?
A pilot tests whether participants understand and follow the recording protocol as written, and whether the equipment performs as expected. Identifying gaps with a small group is inexpensive. Finding the same gaps after hundreds of hours of footage have been collected is not.
