About the Client
Our client is involved in an innovative data-collection project, which focuses
on creating video datasets for training AI models. These AI models are
designed to learn from first-person (POV) video footage. The captured videos
depict real people involved in hands-on tasks with a mounted phone on their
head or chest, paired with spoken narration. The purpose is to enable the
model to connect visual inputs (objects, hands, actions) with audio
descriptions.
Sub-areas for the videos
- Home & Daily Tasks: cleaning; laundry (sorting, folding); organizing a closet; house tours; pet care; packing luggage; loading appliances / dishwashing
- Repairs & DIY: home repair; gardening / farming; furniture assembly; woodworking; plumbing; electrical work; bicycle maintenance; soldering electronics
- Textile & Craft Arts: sewing; knitting; crocheting; using a loom; leathercrafting; bookbinding; crafting jewelry; pottery / ceramics
- Professional Trades & Industrial: automotive repair / maintenance; warehousing / logistics; construction / woodworking; laboratory work; operating heavy machinery controls; assembly-line packaging
- Hobbies & Arts: art – drawing; art – ceramics; art – painting; playing instruments; model building; outdoor survival / camping; calligraphy
- Technology & Computing: using computers or devices (e.g. audio mixers); gaming; product demos; VR / AR interaction; 3D printer setup / maintenance
- Outdoors & Activity: city tours; shopping; navigating public transit
- Personal Care: haircut; applying makeup; detailed grooming routines
- Specialized & Professional: medical procedures; first aid training; professional barista workflows; culinary chef / knife work; lab protocols (pipetting, titrations)
About the Role - Narrator
Watches the POV video and describes the scene and task in first person, in
English, as if explaining it to a model that can't see. Fluency and describing
ability matter most; knowing the activity is not required.
- Deliverable — audio narration synced to the video, recorded after the footage
- Language — English; fluency is criterion #1
- Density — at least 25 words/min on average, no long silences
- Content — 90%+ about the setting or task steps
Required Skills
- English fluency
- Ability to narrate pre-recorded videos
- All details on the Acceptance Criteria will be shared prior to the work.
Rules
- No personal data on video — third-party faces, screens, documents, plates, addresses, mirror reflections.
- 100% human narration — no TTS or synthetic voice.
Engagement Details
- Compensation model: Hourly paid, considering the total number of hours for approved videos.
- Commitment Type : Flexible hours based on video submissions.
- Duration : Ongoing, with earnings based on the number of approved videos.
- Location : Remote, with flexible overlap with client's timezone.
- Start Date : Immediate start upon approval of test video.