You Can't Scrape a Grip
Robot hands can finally feel. Almost nobody has the data to teach them what they're feeling.
For most of the last decade, the honest answer to "why can't robots handle real objects?" was the hardware. The hands were too stiff, too slow, too blind to contact. You could write a beautiful policy and still watch a gripper crush an egg or drop a wet plate, because nothing in the system knew the difference between holding and squeezing.
That answer expired in 2026. Look at what shipped this year. In July, 1X unveiled a new tendon-driven hand for its NEO humanoid: 25 degrees of freedom, with 22 actuated in the fingers and palm and three at the wrist, plus high-resolution tactile sensing across the fingertips and contact surfaces. The sensors read normal force, contact location, and shear, which is the sideways force that tells a hand something is starting to slip. Slip detection happens in real time, so the grip corrects before the glass hits the floor. The company says it has a line capable of building roughly 10,000 hands this year. 1X is not alone. XELA Robotics has been pushing its uSkin tactile technology across a much larger surface area than fingertips alone, covering phalanges and palm, with sensitivity down to 0.1 gram-force, and has integrated it into Tesollo's five-fingered DG-5F hand. Unitree's tactile Dex5-1P variant reportedly adds 94 sensors. CMU's LEAP Hand attacks the same problem from the research end at under $2,000 a unit. At ICRA and Automate this year, the tactile section of the floor was crowded: magnetic, piezoresistive, optical, visuo-tactile, force-inferred. No convergence on modality yet, but no shortage of hardware either. So here is where that leaves every team building manipulation policies in 2026. The bottleneck moved. It isn't the hand anymore. It's the contact data. This distinction gets flattened constantly, and it costs teams entire quarters. Vision data has a shortcut. The internet is a hundred-billion-image corpus that somebody else already paid to create. You can pretrain on it, fine-tune on it, and get remarkably far before you ever put a camera on a robot. Touch has no such corpus. There is no archive of grips. There is no scrape for friction. Every usable hour of tactile data has to be physically produced. A human being has to touch a real object, in a real environment, with a real sensor recording, and then do it again across every variation that matters: wet versus dry, full versus empty, rigid versus deformable, smooth versus textured, one-handed versus two, and the dozen ways the same task fails. That last category is the one teams underweight. A policy trained only on successful grasps has never seen the thing it most needs to recognize, which is the half-second before a failure. Slip, deflection, unexpected compliance, an object that turns out to be heavier than it looks. Those events are rare by definition and cannot be manufactured on demand in a lab. They show up in volume, in real environments, or they don't show up at all. Simulation is not the enemy here. Anyone selling you "sim is fake" is selling you something. Sim gets you into the neighborhood fast. It is cheap, parallelizable, and genuinely excellent for locomotion, navigation, and coarse manipulation. Domain randomization has gotten good. But contact physics is precisely where simulators are weakest. Deformables, friction coefficients, compliance, slip onset, surface variation. These are the hardest things to model accurately and the easiest things to get subtly wrong in ways that only surface as a failure rate on real hardware three months later. The gap does not announce itself. The policy just quietly underperforms in the field. The honest framing is this: synthetic data is a multiplier on real data, not a substitute for it. It drifts, and it needs real-world data to stay anchored. The teams getting the most out of simulation in 2026 are the ones with the largest, most diverse real-contact corpus to calibrate against, not the ones trying to avoid collecting one. When a team says they need tactile data at scale, they usually mean four things at once, and only one of them is volume. Volume. Enough trajectories per task that the tail behaviors appear, not just the median. Diversity. Different homes, different lines, different objects, different operators, different hands. A thousand hours in one kitchen is one kitchen, sampled a thousand times. Specification. Data collected to your task list, your sensor suite, your annotation schema, your file format. Not a general-purpose dump you spend two months reshaping. Provenance. You know who collected it, where, under what consent, and under what license. This matters more every quarter as commercial deployments face procurement review and customers start asking where the training data came from. Volume without the other three is a storage bill. Fizzion collects in two environments, deliberately. Residential, at scale. Homes are the hardest manipulation environment that exists, which is exactly why they are the most valuable place to collect. Laundry, dishes, cords, cabinet handles, groceries, half-full containers, soft goods that change shape when you touch them. Nothing is fixtured. Nothing repeats exactly. No lab can stage this convincingly, and no two homes generate the same data, which is the whole point. Factories and small manufacturers, 5-20+ employees. The other half of the picture. Repeatable fine manipulation, tool use, part handling, real line conditions, real tolerances, real throughput pressure. Structured enough to control, real enough to matter. And critically, small and mid-size operations are underrepresented in existing datasets, which are heavily skewed toward a handful of large flagship plants. Either one. Or both. Most teams need both eventually, because a policy that only ever saw structured environments generalizes badly to unstructured ones, and a policy that only ever saw chaos never learns precision. This is the part we care most about. Fizzion operates a network of 600,000+ collectors. We are not licensing someone else's corpus and marking it up. We are not aggregating scraped video of uncertain origin. We go and collect, to your spec, in the environment you name, with the modalities you need. That distinction shows up in three practical ways. You can request specific tasks rather than searching for them in a dump. You can get the failure cases, because we can instruct for them. And you know the provenance, because it starts with us. The market is starting to treat contact data as infrastructure rather than a research nicety. In Japan, Kawasaki, FANUC, and Yaskawa were selected under METI and NEDO's GENIAC program to jointly build a dataset for VTLA models, integrating vision, touch, language, and action, over a twelve-month program running from this month through July 2027. China's MIIT has launched a national real-world training campaign for embodied intelligence, explicitly citing fragmented, non-interoperable data collection as the thing holding the industry back from economies of scale. Read that as a signal. When national industrial policy starts funding touch datasets, the people building products should assume the data layer, not the hardware layer, is where the next two years of competitive separation happens. The hands are arriving. Ten thousand of them from one vendor alone, this year. The sensors are getting cheaper, denser, and more sensitive every quarter. The question in front of every manipulation team right now is not whether their robot can feel. It's whether they have taught it what it's feeling, across enough of the real world that the lesson holds. That part we can help with. If you're training a model or a robot on tactile data at scale, let's talk. https://calendly.com/robotics-teleoperations-and-end-to-end-training/30min?month=2026-08 Fizzion is a first-party robotics training data company. We collect multimodal human demonstration data at scale for physical AI and autonomous systems. Source, not reseller.The hardware question is closing
Tactile data is not vision data
Where simulation drifts
What "at scale" actually means
The two places contact happens all day
Source, not reseller
One thing worth watching
If touch is on your roadmap
Did you enjoy this article?