NOLO EgoCapture — First-Person Multimodal Data Capture Device

Built for Embodied AI Model Training

NOLO EgoCapture First-Person Multimodal
Data Capture Device

Split design · head unit approx. 210g. Just wear it and capture imperceptibly in real environments — factories, homes, stores and more. Vision, depth, pose and tactile data: five synchronized multimodal streams, ready for model training.

NOLO EgoCapture head-mounted device
NOLO EgoCapture head unit

Real human work is data production.
Moving data capture from “purpose-built setups” to “something that simply happens”.

Other ego datasets are video only. Every frame NOLO EgoCapture outputs carries millimeter-level 6-DoF pose ground truth. A worker assembling parts on the shop floor, a cook working in a kitchen, a clerk stocking shelves — while they do their normal jobs, vision, depth, hand motion and tactile information are recorded in sync and turned directly into high-quality robot training data.

01

Data Is Expensive

Traditional capture depends on the robot itself and on purpose-built capture environments. Hardware and deployment cost a great deal, and the cost per usable sample stays high.

02

Data Volume Is Small

Calibrating and capturing person by person, with limited capture rooms and workstations, makes it hard to reach the tens of millions of samples training requires.

03

Real Scenarios Are Scarce

Capture environments are disconnected from real production, scenario diversity is insufficient, model generalization suffers, and the “last mile” never gets closed.

NOLO SOLUTION

A first-person capture system that needs no robot

The EgoCapture headset and the OmniGlove smart glove form a lightweight capture kit: no robot body to depend on, no complex deployment — put it on and start collecting. Powered by our in-house PolarTraq® and StarTraq® positioning technologies, it holds sub-millimeter accuracy and millisecond latency even in complex real environments, outputting five synchronized multimodal streams at 30Hz, ready for model training.

  • Real people “work as they capture” — data is generated naturally in real settings
  • Every frame carries millimeter-level 6-DoF pose ground truth — trainable out of the box
  • Supports 1,000+ concurrent captures and uploads; 500GB per person per 8 hours
SPLIT DESIGN

Split design — the head unit only captures

The battery and storage move down into the handheld controller, keeping the head unit extremely light and comfortable over long sessions — capturing becomes truly imperceptible.

EgoCapture head unit
HEAD UNIT

Approx. 210g lightweight headset

Only what capture requires: six cameras, a high-precision IMU and dual noise-cancelling microphones. With battery and storage moved elsewhere, the head unit stays extremely light and comfortable through long sessions.

Six-camera ultra-wide FOVIMU 1000HzDual-mic noise cancelling
USB-C LINK
Handheld controller
HANDHELD CONTROLLER

A 22000mAh battery and data hub

Battery, storage and capture control are centralized in the handheld terminal. A 4.5-inch touchscreen handles live preview and task management, and 40W fast charging supports about 8 hours of continuous capture.

22000mAh battery512G storageTouch preview
CORE CAPABILITIES

Put it on and go — six core capabilities

From ultra-wide field-of-view capture to ground-truth-grade data output: a complete multimodal data infrastructure for embodied AI model training.

01

Six-Camera Ultra-Wide FOV

Five global-shutter RGB cameras plus one ToF depth camera deliver a full-device image FOV of 208°×193°, giving all-around visual coverage from the head-worn viewpoint, with depth and color captured in sync.

Per camera 1920×1200 @ 60Hz | ToF 640×480 @ 30Hz
02

Millimeter-Level 6DoF Ground Truth

Head, wrist and finger poses are output per frame with millimeter-level accuracy, strictly time-aligned with the video frames. The data is ready for motion retargeting and model training, removing tedious SLAM post-processing and failed re-captures.

6DoF accuracy 10mm | strictly synchronized with video
03

Five Modalities Synchronized at 30Hz

Five data streams — RGB imagery, ToF depth, audio, gesture and tactile — are output at 30Hz with unified timestamps, keeping multi-device capture spatiotemporally consistent and fit for large-scale dataset construction.

Vision / depth / audio / gesture / tactile, synchronized
04

Lightweight Split Headset

Battery and storage move down to the handheld controller; the head unit keeps only the capture module. At roughly 210g it stays comfortable through long shifts for truly imperceptible capture.

Head unit approx. 210g | split design
05

8-Hour Battery Life

The handheld controller houses a large 22000mAh battery and supports 40W PD fast charging, delivering around 8 hours of continuous capture per charge for full-day operation and long task capture.

22000mAh | 40W fast charging | approx. 8 hours
06

Multi-Device Spatiotemporal Alignment

Supports 1,000+ people capturing and uploading at the same time, with data solving capacity matched 1:1 to capture capacity. Timestamps across streams stay strictly aligned when multiple people and rigs capture together, scaling up training datasets.

1,000+ concurrent | 1:1 data solving
COLLECTION MODES

Two capture modes, from fine manipulation to dexterous work

The same EgoCapture headset, paired with the OmniGlove smart glove or with bare-hand tracking, switches capture granularity to fit the task.

FINE-SKILL DEMONSTRATION NOLO OmniGlove smart glove

OmniGlove Smart Glove

NOLO OMNI GLOVE

Captures fingertip tactile feedback and dexterous manipulation with high precision: the tactile array outputs per-finger normal force values and tangential force directions in real time. Paired with the headset for fine-skill demonstration, it suits assembly, caregiving and surgery — anywhere dexterity matters.

Five-finger tactile array Normal force + tangential direction Dexterous hand motion capture
WEAR AND GO NOLO EgoCapture bare-hand tracking

Bare-Hand Tracking

BARE-HAND TRACKING

No glove required: the headset's built-in algorithm identifies 21 hand keypoints in real time. Raise your hand and it captures, with natural unrestricted motion — ideal for batch data capture of high-frequency daily tasks such as retail restocking and home cleaning.

21 hand keypoints Real-time recognition Zero wearables
DATA CAPABILITIES

From capture to ground truth — one complete data pipeline

Capture end, sensing layer, algorithm layer and output layer are connected end to end, delivering multimodal datasets that go straight into model training.

CAPTURE END
  • EgoCapture headset
  • OmniGlove smart glove / bare hand
  • Real people working in real scenarios
SENSING LAYER
  • 5× global-shutter RGB cameras
  • 1× ToF depth camera
  • IMU @1000Hz | dual-mic noise cancelling
  • Tactile array (OmniGlove)
ALGORITHM LAYER
  • StarTraq® positioning: millimeter-level 6DoF for head / wrist / fingers
  • Gesture recognition: bare-hand + glove dual mode, 21 keypoints
  • Tactile solving: normal force values + tangential force direction
  • Multi-device spatiotemporal alignment with unified global timestamps
OUTPUT LAYER
  • 5-channel RGB video + ToF depth maps
  • Head / wrist 6DoF pose ground truth
  • Hand 21-keypoint tracking
  • Tactile data (force / direction)
  • Five modalities synchronized at 30Hz
1000+
1,000+ concurrent capture

Supports more than a thousand people capturing and uploading at once, meeting the throughput demands of a scaled data factory.

500G
500GB per person per 8 hours

A single full capture session yields hundreds of gigabytes of raw data, covering an entire day of work.

1:1
1:1 capture-to-solving

Data solving capacity matches capture capacity 1:1 — solving completes as capture finishes, with no queueing.

APPLICATION SCENARIOS

Step into real scenarios — data happens naturally

For embodied AI robot training in industrial manufacturing, domestic services, retail and medical care — real people “work as they capture”.

Industrial manufacturing scenario
SCENE 01

Industrial Manufacturing

Capturing fine operations in shop-floor assembly, machine tending and quality inspection

Domestic services scenario
SCENE 02

Domestic Services

Capturing everyday household motions such as home cleaning and tidying up

Retail and supermarket scenario
SCENE 03

Retail & Supermarkets

Capturing store operations such as shelf restocking and item picking

Medical care scenario
SCENE 04

Medical Care

Capturing clinical skills such as nursing procedures and instrument handling

SPECIFICATIONS

Core Specifications

Key parameters of the EgoCapture headset and the handheld controller. Final values are subject to the delivered version.

EgoCapture Headset

EGO HEAD UNIT
ProcessorRK3588, local or real-time upload; offline output of 6DoF and 21 hand keypoints
Memory / StorageLPDDR5 12GB | UFS3.1 512G | TF card 512G
RGB Cameras5× global shutter, per camera 1920×1200 @60Hz, full-device image FOV 208°×193°
Depth Camera1× ToF, 640×480 @30Hz, FOV 121.5°×98°
PositioningFull-device positioning FOV 180°×160°, 6DoF accuracy 10mm
IMUAccelerometer + gyroscope, sampling at 1000Hz
IR Illumination5× 850nm LED, FOV 150°; visual enhancement for more robust positioning
Microphones2× with noise reduction
WirelessWiFi 6 | 2.4G private link (IR positioning module expansion)
PortsType-C ×2 (handheld terminal link / USB3.1 wired real-time upload)
WeightHead unit approx. 210g (soft head strap excluded)
Split DesignBattery, storage and capture control sit in the handheld controller; the head unit keeps only the capture module
SystemCustom capture system based on Android 14

Handheld Controller

HANDHELD CONTROLLER
Display4.5-inch touchscreen, resolution 450×845
Battery22000mAh, 40W PD fast charging, approx. 8 hours of runtime
PortsType-C ×2 (headset link / charging)
ButtonsPower button ×1
Companion AppDAS app (Android / iOS) for task management, data capture, preview and upload
Capture ControlCapture control button, capture status indicator

* Specifications are taken from official NOLO product documentation and may change as versions are updated. Please refer to the delivered product.

Turn real work into training data for robots

Book a product demo, or discuss your embodied AI data capture plan with us. Based on your scenario, the NOLO team will advise on everything from hardware selection to closing the data loop.

Business & Partnership