Robot data collection cost is the total spend required to produce a usable demonstration dataset, and it lands between roughly $2 and $150 per usable demonstration depending on the task. This guide prices a fixed deliverable of 1,000 demonstrations three ways. It breaks a quote into the seven cost lines sitting underneath it, and shows why a first program costs about twice what the second one does. You also get the arithmetic that decides whether the spend turns into policy performance.
Most cost guides answer with an hourly rate. An hourly rate tells you almost nothing, because two teams paying the same rate for the same task can end up 5× apart on cost per usable demonstration. The gap sits in throughput, acceptance rate, and the one-time engineering that never appears on a collection quote.
The model below uses 1,000 demonstrations as the unit because it is a realistic first commercial program: large enough to train a single-task policy that generalises, small enough to fund before the results are proven.
What Does Robot Data Collection Cost per Demonstration?
Cost per usable demonstration is the number to quote and the number to negotiate. It combines three variables that vendors usually report separately: the all-in operator hour, how many demonstrations that hour produces, and what share of them survive quality review.
The formula is simple, and writing it out changes how a quote reads:
Cost per usable demonstration = all-in hourly rate ÷ (recorded demonstrations per hour × acceptance rate)
| Task tier | Recorded demos per operator hour | Typical acceptance | All-in operator hour | Cost per usable demonstration |
|---|---|---|---|---|
| Single-arm pick, place and sorting | 10–15 | 85–92% | $25–40 | $2–5 |
| Bimanual fine manipulation | 3–6 | 75–85% | $35–60 | $7–27 |
| Contact-rich work with tactile and force capture | 2–4 | 65–80% | $40–70 | $13–54 |
| Mobile manipulation and humanoid whole-body | 1–3 | 60–75% | $45–90 | $20–150 |
These figures cover collection labour only. Hardware, pipeline engineering and program management sit on top, and the next two sections price them.
7 Cost Lines Behind a Robot Data Collection Quote
A quote usually shows one or two lines. A program has seven. Teams that budget from the quote miss the difference and run out of money at about the 60 percent mark.
| Cost line | What it covers | Share of a contracted program |
|---|---|---|
| Operator time | Recording, scene resets, object staging, retakes | 55–70% |
| Supervision and QA review | Episode grading, sampling, rejection handling | 8–15% |
| Pipeline and format conversion | Segmentation, sync checks, LeRobot or HDF5 packaging | 5–15% |
| Program management and spec iteration | Task definition, revisions, weekly sample reviews | 8–12% |
| Rig hardware amortisation | Arms, cameras, sensors, servicing, spares | 3–10% contracted, far higher in-house |
| Consent, provenance and licensing admin | Operator releases, per-episode chain of custody, rights terms | 1–3% |
| Storage, transfer and archive | Object storage, egress, cold copies | Under 2% |
Treat those shares as planning ranges rather than a balanced budget. Two of them deserve attention up front. Program management is where a vague task specification becomes expensive, because every revision invalidates episodes already recorded. Consent and provenance is small as a percentage and decisive as a risk: a dataset without signed releases and a per-episode record of who produced it on what hardware is difficult to license, sell or defend later.
Our physical AI training data collection programs treat that record as part of the deliverable rather than paperwork attached to it.
Pricing 1,000 Demonstrations: 3 Worked Budgets
Below are three complete budgets for the same deliverable of 1,000 usable demonstrations. Each uses the throughput and acceptance figures from the tier table, and each states its assumptions so you can substitute your own.
| Budget line | A. Single-arm, contracted | B. Bimanual, contracted | C. Tactile, built in-house |
|---|---|---|---|
| Acceptance rate assumed | 88% | 78% | 70% |
| Episodes that must be recorded | 1,140 | 1,280 | 1,430 |
| Operator hours required | 95 | 320 | 480 |
| Collection labour | $2,850 | $15,360 | $21,600 |
| Rig hardware and sensors | In rate | In rate | $26,000 |
| Rig build and operator training | In rate | In rate | $12,000 |
| Pipeline and format tooling | $1,200 | $2,000 | $8,000 |
| QA review and curation | In rate | In rate | $6,000 |
| Program management | $340 | $1,840 | $9,000 |
| Storage and transfer | $40 | $110 | $400 |
| Program total | $4,430 | $19,310 | $83,000 |
| Cost per usable demonstration | $4.43 | $19.31 | $83.00 |
Assumptions: recorded throughput of 12, 4 and 3 demonstrations per operator hour; all-in operator rates of $30, $48 and $45; engineering time costed at $100 per hour; program management at 12 percent of collection labour for the contracted routes. Substitute your own rates and the structure holds.
The spread across those three columns is nearly 19×, and every column is defensible. Cost is set by what the task demands, so a quote that looks expensive against a benchmark from a different tier is being compared against the wrong thing.
A note on the figures. Every rate and total in this guide is a planning model rather than a quoted price. Cost per usable demonstration moves with task complexity, sensor coverage, operator skill and tenure, rig type, acceptance thresholds, the scene and object diversity you specify, program duration, location and labour market, and how much hardware and pipeline you already own. Two vendors quoting an identical hourly rate can finish far apart on total cost once throughput and acceptance are applied. Use these numbers to structure a budget and to compare quotes on the same terms, then price your own program against your task specification and a partner’s measured yield.
Why the First 1,000 Demonstrations Cost More Than the Next 1,000 Demos?
Column C carries $46,000 of spend that happens once: hardware capital, rig build, operator training, and the pipeline that turns raw recordings into training-ready episodes. Run a second program of 1,000 demonstrations on the same rig with the same trained operators and those lines drop out.
| In-house program | Total cost | Cost per usable demonstration |
|---|---|---|
| First 1,000 demonstrations | $83,000 | $83.00 |
| Second 1,000 on the same rig | $37,000 | $37.00 |
| Blended across both | $120,000 | $60.00 |
That 2.2× ratio decides build against buy more cleanly than any hourly comparison. Building your own floor pays off when you will collect continuously for many months on tasks specific to your product. Contracting pays off for a bounded program, for scene breadth, and for the first program of all, where the honest question is whether the data moves policy success at all. Teams often answer that question with a contracted pilot and build the floor afterwards, once the tasks worth owning are known.
Does Offshore Robot Data Collection Cost Less?
Wages differ by a large multiple. Program cost differs by much less, and the reason is visible once the rate stack is separated by layer. Reporting from HelloChinaTech on Chinese collection facilities puts an operator’s wage at RMB 20 to 40 per hour, while a data platform with a robot in the loop quotes RMB 500 to 1,000 per hour for the same hour of output. Galaxea’s chief executive put internal production cost near RMB 50 to 100 per hour for wearable capture and around RMB 250 with a robot involved.
| Layer | Reported rate per hour | Roughly, in USD | What sits inside it |
|---|---|---|---|
| Operator wage | RMB 20–40 | $3–6 | The person only |
| Production cost, wearable capture | RMB 50–100 | $7–14 | Operators, rigs, floor space |
| Production cost, robot in the loop | About RMB 250 | About $35 | Adds robot depreciation and maintenance |
| Commercial service rate, with robot | RMB 500–1,000 | $70–140 | Rejected episodes, QA, margin |
Converted at roughly RMB 7.1 to the dollar, the top layer sits above typical North American all-in benchmarks of $28 to $60 per operator hour. Wage arbitrage shrinks as soon as a robot enters the loop, because depreciation, floor space and rejected episodes carry the same cost anywhere. The saving that survives is on wearable and robot-free capture, where labour is the dominant input.
Two questions settle it faster than a rate comparison.
- Ask what the quoted hour includes, layer by layer.
- Then ask for usable episodes per operator hour, because a $30 hour at 55 percent yield loses to a $50 hour at 90 percent yield.
What 1,000 Demonstrations Buys, and How to Spend Them?
A budget for 1,000 demonstrations is also a decision about where those demonstrations get recorded, and that decision changes policy performance more than the price per episode does.
The Data Scaling Laws in Imitation Learning study collected over 40,000 demonstrations and ran more than 15,000 real-world rollouts to test the question directly. Generalisation scaled as a rough power law with the number of environments and objects, while demonstrations per environment hit a threshold beyond which additional ones changed little. The team trained policies reaching about 90 percent success on unseen environments and objects using 32 environment-object pairs at 50 demonstrations each, a total of 1,600.
Applied to a 1,000-demonstration budget, that points to roughly 20 environment-object pairs at 50 demonstrations each rather than 1,000 repetitions on one bench. The second allocation costs the same and generalises far less. It also changes what to ask a collection partner for: scenes, lighting conditions, object instances and viewpoints, specified explicitly, rather than a single episode count.
Expect rejection to be normal rather than a supplier failure. The DROID dataset recorded 76,000 successful trajectories and labelled roughly 16,000 more as unsuccessful, which is a rejection rate near 17 percent inside a carefully run academic program with a standardised hardware stack. A commercial floor working on a newly specified task will do worse in week one and better by week four.
Storage, Egress and the Costs Teams Overestimate
Storage gets listed as a hidden cost in most vendor material. At 1,000 demonstrations it is noise, and the arithmetic shows why.
DROID reports 350 hours of interaction across 76,000 episodes, which works out to roughly 17 seconds of recorded interaction per episode. A 20-second episode carrying three compressed RGB streams plus proprioception runs about 80 to 150 MB; add stereo depth and tactile channels and it reaches 300 to 500 MB. One thousand episodes therefore occupy somewhere between 0.1 and 0.5 TB.
At AWS S3 Standard pricing of $0.023 per GB-month in US East, 300 GB costs about $7 per month to store and about $27 to transfer out once at $0.09 per GB. Those are rounding errors against a five-figure collection budget.
Storage becomes a real line at a different scale. A 100,000-episode corpus at 300 MB per episode is 30 TB, around $690 per month in standard storage, and every full copy pulled out costs roughly $2,700 in egress alone. Plan the storage class and the copy policy at that point, not before.
The cost teams underestimate instead is spec iteration. Change the object set, the camera placement or the success criterion after week two, and the episodes recorded before the change may no longer be trainable alongside the ones recorded after. That is a re-collection cost, and it is charged at the full rate.
Scoping a robot data collection budget or comparing quotes across vendors? NeuralChainAI runs robot data collection services across teleoperation and egocentric capture, including tactile manipulation data with force arrays, consented at source with a per-episode chain of custody. Bring the task list and the target policy; you get back a costed plan with throughput, acceptance and allocation written down before episode one.
What to Demand on a Robot Data Collection Quote?
The following items turn a rate into a budget you can hold a vendor to. Ask for all of them in writing.
- Usable episodes per operator hour, separated from recorded episodes per hour. One number is throughput, the other is what you pay for.
- The acceptance rate and the rejection criteria that produce it, plus who grades episodes and at what sampling rate.
- Scene and object allocation: how many distinct environments, object instances, lighting conditions and viewpoints the episode count is spread across.
- Delivery format and schema, named specifically. LeRobot and HDF5 are the common targets, and retrofitting a format onto recorded episodes costs engineering time you have not budgeted.
- Sensor coverage per episode, listed channel by channel with sample rates, including which streams are hardware-synchronised rather than aligned afterwards.
- Provenance and consent terms: signed operator releases, a per-episode record tying data to the person and hardware that produced it, and explicit rights to train, deploy and redistribute.
- Re-collection terms for spec changes, including who absorbs the cost when the task definition moves.
A vendor who answers all these with numbers is quoting a program. A vendor who answers with an hourly rate is quoting an input. Teams weighing a data budget alongside a wider build can scope both together through physical AI and robotics consulting, where hardware, capture and model strategy get costed as one plan.