Headquarters:
Our client is involved in an innovative data-collection project, which focuses on creating video datasets for training AI models. These AI models are designed to learn from first-person (POV) video footage. The captured videos depict real people involved in hands-on tasks with a mounted phone on their head or chest, paired with spoken narration. The purpose is to enable the model to connect visual inputs (objects, hands, actions) with audio descriptions.
Sub-areas for the videos
Home & Daily Tasks:
cleaning; laundry (sorting, folding); organizing a closet; house tours; pet care; packing luggage; loading appliances / dishwashing
Repairs & DIY:
home repair; gardening / farming; furniture assembly; woodworking; plumbing; electrical work; bicycle maintenance; soldering electronics
Textile & Craft Arts:
sewing; knitting; crocheting; using a loom; leathercrafting; bookbinding; crafting jewelry; pottery / ceramics
Professional Trades & Industrial:
automotive repair / maintenance; warehousing / logistics; construction / woodworking; laboratory work; operating heavy machinery controls; assembly-line packaging
Hobbies & Arts:
art – drawing; art – ceramics; art – painting; playing instruments; model building; outdoor survival / camping; calligraphy
Technology & Computing:
using computers or devices (e.g. audio mixers); gaming; product demos; VR / AR interaction; 3D printer setup / maintenance
Outdoors & Activity:
city tours; shopping; navigating public transit
Personal Care:
haircut; applying makeup; detailed grooming routines
Specialized & Professional:
medical procedures; first aid training; professional barista workflows; culinary chef / knife work; lab protocols (pipetting, titrations)
Watches the POV video and describes the scene and task in first person, in English, as if explaining it to a model that can't see. Fluency and describing ability matter most; knowing the activity is not required.
Deliverable
— audio narration synced to the video, recorded after the footage
Language
— English; fluency is criterion #1
Density
— at least 25 words/min on average, no long silences
Content
— 90%+ about the setting or task steps
Required Skills
English fluency
Ability to narrate pre-recorded videos
All details on the Acceptance Criteria will be shared prior to the work.
Rules
No personal data on video — third-party faces, screens, documents, plates, addresses, mirror reflections.
100% human narration — no TTS or synthetic voice.
Engagement Details
Compensation model:
Hourly paid, considering the total number of hours for approved videos.
Commitment Type
: Flexible hours based on video submissions.
Duration
: Ongoing, with earnings based on the number of approved videos.
Location
: Remote, with flexible overlap with client's timezone.
Start Date
: Immediate start upon approval of test video.
To apply:
https://weworkremotely.com/remote-jobs/toptal-narrator-videos-for-data-collection-project-ai-training
Browse by category