Dataset & Benchmark

SoftVTBench

A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation

Touch tells the policy how contact is evolving. FEM tells the evaluator what that contact did to the object.

Bowen Jing1,*, Mingxin Wang1,2,*, Ruiyang Hao3, Chenchen Ge1,4, Hanwen Shen5, Junjie He6, Yang Cui7, Yiming Hou1,4, Weitao Zhou2,8,‡, Jiawei Wang8, Minglei Li8, Dandan Zhang9, Ding Zhao10, Houde Liu2, Xiaofan Li11, Si Liu12, Ping Luo13, Haibao Yu1,13,‡

1Tuojing Intelligence  ·  2Tsinghua University  ·  3King's College London  ·  4Southeast University  ·  5Stevens Institute of Technology  ·  6The Hong Kong University of Science and Technology (GZ)  ·  7University of Manchester  ·  8Simple AI  ·  9Imperial College London  ·  10Carnegie Mellon University  ·  11Zhejiang University  ·  12Beihang University  ·  13The University of Hong Kong

* Equal contribution    ‡ Corresponding author

🎞️ Expert Demonstrations 4,000
🧩 Tasks / Suites 40 / 4
🧈 Assets 50+
🤖 Streams @ 20 Hz Multi-View RGB + Tactile RGB + Marker Motion + Proprioception + Language + Actions
Overview of SoftVTBench: four diagnostic suites, 4,000 demonstrations, 50+ assets including deformable objects and matched rigid twins, controlled visual and physical shifts, the interaction regimes from slip-prone loose grasps to excessive compression, and the synchronized visual, tactile, proprioceptive, language, and action streams
4,000 demonstrations, four diagnostic suites, 50+ assets — deformable objects and their visually matched rigid twins. Between a grasp too loose to hold and one tight enough to crush lies a narrow safe window, and only touch observes it from the inside.

Abstract

A policy can complete a manipulation task while letting the object slip — or by crushing it. Task success sees neither. SoftVTBench is a visuo-tactile dataset and benchmark that separates the two. It contains 4,000 expert demonstrations over 40 tasks and 50+ assets, each episode synchronizing multi-view RGB, dual-finger tactile RGB and marker motion, proprioception, language, and actions at 20 Hz — alongside evaluator-only finite-element (FEM) states the policy never sees. On top of it we define the Deformation-aware Success Rate (DSR), which credits a rollout only when the task is completed and deformation stays within a per-object tolerance calibrated before any policy is trained. Across Diffusion Policy, π0.5, and FastWAM, every one of the 12 in-distribution configurations contains successes that violate that tolerance. Under distribution shift, visuo-tactile variants win all six task-success comparisons and five of six on DSR — while in distribution the same comparison is split. Making touch available is not the same as using it.

The Metric

Task success is a component of DSR, not an alternative to it.

DSR = Task Completed and Deformation Within Tolerance

The task predicate is purely kinematic: the instructed object comes to rest in the target region. The deformation side is read from FEM states the policy never sees, and compared against a tolerance calibrated per object before any policy is trained — so no method can move the bar it is scored against. Reporting Task Success Rate (TSR) alongside DSR makes their difference an exact count of the rollouts that reached the target by mishandling the object.

What the policy sees

Third-person and wrist RGB, dual-finger tactile RGB and marker motion, proprioception, and the language instruction. It outputs an end-effector pose target plus a gripper command — binary or continuous.

What the evaluator sees

FEM nodal positions, object poses, contact and drop events. Deformation is the peak rigid-motion-removed RMS node displacement, normalized by object size — which separates being carried from being squeezed.

Calibrated up front

A scripted grasp–lift–hold sweep fixes each object's gripper envelope and its deformation tolerance before training. The full trace is released, so the criterion can be audited or rescored without new rollouts.

SoftVTBench construction and evaluation pipeline: stage 1 matched rigid-deformable object construction and per-object grasp calibration, stage 2 controlled task generation and constraint-aware expert rollout with policy-visible observations separated from policy-hidden states, stage 3 integrity checking, outcome labeling into TSR and DSR, human verification, and the released train / ID / OOD splits
Three stages: build matched rigid–deformable objects and calibrate their constraints; generate controlled tasks and record policy-visible observations separately from evaluator-only states; screen, verify, and label into the released train / ID / OOD splits.

Where SoftVTBench Sits

Complete tasks, volumetric deformable objects, policy-visible touch, physical ground truth, deformation-aware evaluation — existing suites have some, none have all five.

Benchmark Complete Task 3D Deformable Policy-Visible Touch Physical Ground Truth Deformation-Aware Eval
LIBERO
ManiSkill2
SoftGym
MoDeSuite
DefGraspSim
SoGraB
VTDexManip
ManiFeel
Tabero
SoftVTBench

✓ full support  ·  ◎ partial  ·  ✗ not supported.

Dataset

Every episode records two views of the same interaction: what the policy may use, and what only the evaluator may read.

4,000 Expert Demonstrations
40 Tasks in 4 Suites
50+ Assets
20 Hz Synchronized Streams
SuiteObject TypeVariation Axis#Tasks#DemosID Eval EpisodesOOD Conditions
Object-SoftDeformableObject identity101,0005009
Spatial-SoftDeformableSpatial layout101,0005009
Object-RigidRigid twinObject identity101,000500
Spatial-RigidRigid twinSpatial layout101,000500
Total404,0002,000

OOD covers nine held-out conditions on the two deformable suites — three levels each of lighting (67.5 / 180 / 270), mass (×1.25 / ×1.75 / ×2.5), and Young's modulus (×0.5 / ×0.8 / ×2.0). One factor moves at a time; the task, initial state, and seed are reused from its ID reference.

Tactile RGB images with marker-motion overlays at the moment of grasping, for the ten Object-Soft tasks
Tactile RGB with marker-motion overlay at the moment of grasp, for the ten Object-Soft objects. Different geometry, compliance, and contact area produce visibly different contact patches and shear fields — cues no external camera provides.
Policy-visiblethird-person + wrist RGB · dual-finger tactile RGB + 11×9 marker fields · proprioception · language · actions in both binary and continuous gripper encodings
Evaluator-onlyFEM nodal positions · object poses · contact and drop events — released in full, never exposed to a policy
SimulatorIsaac Sim 4.5.0 · Isaac Lab 0.41.3 · PhysX 5 GPU FEM · Franka + Panda gripper · 60 Hz physics, 20 Hz control
TactileGelSight Mini via TacEx · Taxim optics · FOTS marker motion
Quality controlautomatic integrity screening + human verification on every episode
Headroomexpert demos reach a median 43% of the tolerance; none exceeds it

Task Suites

A matched 2×2 over object type (deformable vs. rigid twin) and variation axis (object identity vs. spatial layout) — one shared pick-and-place skill throughout.

Deformable Main Suites

The object must reach its target without slipping out of the grasp or being compressed past its calibrated tolerance.

OS

Object-Soft

Fixed layout, ten different objects — six bakery-style meshes and four procedural primitives. The grasp must adapt to each object's geometry and compliance.

Deformable
SS

Spatial-Soft

Two identical instances per scene; only the instruction says which one. Success checks the identity of the object that actually moved.

Deformable

Matched Rigid Twin Suites

Same mesh, texture, and mass — stiffness raised until deformation is negligible. Same layouts and instructions, so difficulty on a soft suite cannot simply be blamed on deformability.

OR

Object-Rigid

Object-Soft replicated with rigid twins. Isolates object-identity adaptation from deformability.

Rigid Control
SR

Spatial-Rigid

Spatial-Soft replicated with rigid twins. Isolates language-grounded target selection from deformability.

Rigid Control

Results

Diffusion Policy, π0.5, and FastWAM under paired vision-only (VO) and visuo-tactile (VT) inputs, on identical episodes and seeds.

Hidden by task success

12 / 12

ID configurations containing successful rollouts that exceed the deformation tolerance — 0.7–24% of each one's own successes.

Touch under shift

6 / 6  ·  5 / 6

OOD comparisons won by the visuo-tactile variant on TSR, and on DSR. In distribution, the same comparison is split.

Not an inherent tradeoff

≤ 0.4 pt

FastWAM's TSR–DSR gap on both spatial configurations — two episodes out of 500 — while leading the table.

In-Distribution — TSR vs. DSR

ModelInputObject-SoftSpatial-Soft
TSRDSRTSRDSR
Diffusion PolicyVO-C37.433.615.613.4
VT-C40.030.433.025.0
π0.5VO-C41.638.426.022.6
VT-C41.435.027.622.0
FastWAMVO-C62.058.037.036.6
VT-C57.654.456.456.0

DSR is below TSR in all twelve configurations. For Diffusion Policy VT-C the gap is 9.6 points on Object-Soft — 24% of its own successes, an exact count of 48 episodes. It also flips rankings: TSR puts VT-C above VO-C (40.0 vs. 37.4), DSR reverses it (30.4 vs. 33.6). And the gap is family-dependent: 10–24% for Diffusion Policy, 8–20% for π0.5, but only 0.7–6.5% for FastWAM.

Six illustrative rollouts across Object-Soft and Spatial-Soft, each showing synchronized third-person, wrist, and dual tactile observations at approach, peak interaction, and placement, next to the normalized deformation trace; four blue rollouts stay within tolerance and two orange rollouts exceed it despite task success
Every rollout here succeeds at the task. Blue stays within tolerance; orange crosses it during the grasp and stays elevated through transport — then places the object correctly. The RGB views are hard to tell apart. The marker fields and the deformation trace are not.

Deformable vs. Matched Rigid Twin (TSR)

ModelInputObject variationSpatial variation
RigidSoftRigidSoft
Diffusion PolicyVO-C40.037.414.015.6
VT-C35.040.011.033.0
π0.5VO-C60.041.650.426.0
VT-C59.641.454.027.6
FastWAMVO-C64.062.025.037.0
VT-C61.657.630.056.4

Deformability does not cost the same for everyone — and can pay. π0.5 drops 18.4 and 24.4 points going rigid → soft, while FastWAM gains 12.0 on spatial variation. Diffusion Policy is not language-conditioned, so its spatial numbers are a floor set by that limitation, not a measure of spatial generalization.

Sensing or Control Granularity? (π0.5 ablation)

Both gripper encodings are stored for every demonstration, so the two factors can be crossed without recollecting data. B = binary, C = continuous.

ConfigurationObject-SoftSpatial-Soft
TSRDSRTSRDSR
VO-B30.227.234.220.0
VO-C41.638.426.022.6
VT-B41.028.030.021.4
VT-C41.435.027.622.0

A "tactile gain" can be a gripper gain. From VO-B, continuous control alone adds 11.4 TSR points and touch alone adds 10.8 — together they add nothing further. Meanwhile VO-C and VT-B are 0.6 points apart in TSR but 10.4 apart in DSR. Match the action space before crediting touch.

Out-of-Distribution (Δ vs. ID)

ModelInputObject-SoftSpatial-Soft
TSRDSRTSRDSR
Diffusion PolicyVO-C 29.2 (−8.2) 26.6 (−7.0) 11.0 (−4.6) 8.8 (−4.6)
VT-C 31.2 (−8.8) 25.0 (−5.4) 25.2 (−7.8) 17.8 (−7.2)
π0.5VO-C 35.8 (−5.8) 33.2 (−5.2) 24.4 (−1.6) 19.4 (−3.2)
VT-C 41.0 (−0.4) 34.2 (−0.8) 28.4 (+0.8) 23.2 (+1.2)
FastWAMVO-C 54.4 (−7.6) 53.8 (−4.2) 27.8 (−9.2) 27.2 (−9.4)
VT-C 55.8 (−1.8) 55.8 (+1.4) 39.4 (−17.0) 38.8 (−17.2)
Task success rate under out-of-distribution shifts resolved by factor: rows are the Object-Soft and Spatial-Soft suites, column groups are Diffusion Policy, pi-0.5, and FastWAM, each split into illumination, object-mass, and Young's-modulus shifts, comparing vision-only and visuo-tactile curves
The same shifts resolved by factor — lighting, mass, stiffness — for each policy and suite. 100 episodes per point; bands are task-stratified bootstrap 95% CIs, and the pale strip marks the in-distribution reference.

Touch pays off under shift, not everywhere. VT-C wins all six TSR comparisons (sign test p = 0.016) and five of six on DSR. The pattern follows the suite more than the factor: on Spatial-Soft the tactile curve sits at or above vision-only everywhere, while on Object-Soft the two overlap except under mass shifts — touch only observes what happens after contact. Lighting acts before contact and degrades the pathway both modalities share, which is why Diffusion Policy collapses at both illumination extremes regardless of input.

Caveat. The released Diffusion Policy and FastWAM VT variants train at smaller effective batch sizes, so their VO–VT comparisons describe those variants rather than isolate the modality; the π0.5 ablation is batch-matched. Everything here is simulation, and the gap to physical tactile sensors is not characterized in this work.

Citation

@article{jing2026softvtbench,
  title         = {SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and
                   Benchmark for Deformable-Object Manipulation},
  author        = {Jing, Bowen and Wang, Mingxin and Hao, Ruiyang and Ge, Chenchen and
                   Shen, Hanwen and He, Junjie and Cui, Yang and Hou, Yiming and
                   Zhou, Weitao and Wang, Jiawei and Li, Minglei and Zhang, Dandan and
                   Zhao, Ding and Liu, Houde and Li, Xiaofan and Liu, Si and
                   Luo, Ping and Yu, Haibao},
  year          = {2026},
  eprint        = {2607.04234},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2607.04234}
}