Skip to content

perf(rbd): dedupe shared TriMesh uploads in from_rapier - #19

Closed
haixuanTao wants to merge 1 commit into
dimforge:mainfrom
haixuanTao:perf/trimesh-dedupe
Closed

perf(rbd): dedupe shared TriMesh uploads in from_rapier#19
haixuanTao wants to merge 1 commit into
dimforge:mainfrom
haixuanTao:perf/trimesh-dedupe

Conversation

@haixuanTao

Copy link
Copy Markdown
Contributor

What

RbdState::from_rapier now caches the GPU Shape descriptor of each TriMesh by its parry shape data pointer. When the same SharedShape Arc is cloned across environments — the normal way to build batched scenes with shared terrain — the flat BVH + pseudo-normals are serialized into shape_buffers once and every clone reuses the descriptor (which only holds ranges into the shared buffers).

Why

Batched RL training scenes routinely attach one large terrain trimesh to hundreds or thousands of environments. Before this, from_rapier re-uploaded the mesh per env: with 3 terrain strips of a few thousand triangles across 4096 envs we measured hundreds of MB of duplicated vertex/index/BVH data; after, it's O(unique meshes) — three uploads total.

Scope / safety

  • Cache hit requires the same parry data pointer, i.e. an actual SharedShape clone. Structurally-equal-but-distinct meshes are (correctly) not deduped.
  • TriMesh only — every other shape takes the exact previous path, and scenes without shared trimeshes produce byte-identical buffers.
  • Host-only; no shader change.

cargo check clean on nexus_rbd3d and nexus_rbd2d. Running in production on our terrain-curriculum training (4096 envs × 3 shared strips) for two days.

🤖 Generated with Claude Code

https://claude.ai/code/session_01U2n9RqmxTJb8UG5d1Sjw4W

Batched RL environments typically clone one terrain SharedShape across N
envs; from_rapier re-serialized the flat BVH + pseudo-normals N times. The
GPU Shape descriptor only holds ranges into the shared shape_buffers, so
caching it by the parry shape data pointer uploads each unique trimesh
once — vertex/index memory O(unique meshes) instead of O(envs). TriMesh
only; scenes without shared trimeshes produce byte-identical buffers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U2n9RqmxTJb8UG5d1Sjw4W
@sebcrozet

Copy link
Copy Markdown
Member

This will be merged as part of #36

sebcrozet added a commit that referenced this pull request Aug 29, 2026
* fix(rbd): apply the per-batch stride to collider_parent reads in the narrow phase

Replaces #21

Co-Authored-By: Haixuan Xavier Tao <tao.xavier@outlook.com>

* feat: per-environment collision-pair capacity override

Replaces #24

Co-Authored-By: Haixuan Xavier Tao <tao.xavier@outlook.com>

* feat(rbd): make the narrow-phase contact prediction distance configurable

Replaces #28

Co-Authored-By: Haixuan Xavier Tao <tao.xavier@outlook.com>

* fix(rbd): thread the configurable prediction distance through the brute-force broad phase

Completes #28

* feat(python): per-environment MJCF insertion

Replaces #16

Co-Authored-By: Haixuan Xavier Tao <tao.xavier@outlook.com>

* fix(python): drop the duplicated collisions-capacity setter and pass the RbdCoupling to insert_rigid_body_in

Completes #16

* feat(python): per-step MJCF actuator control + multibody state readback

Replaces #12

Co-Authored-By: Haixuan Xavier Tao <tao.xavier@1ms.ai>

* fix(python): gate multibody control/readback on dim3, add the missing PyArray2 import

Replaces #12

Co-Authored-By: Haixuan Xavier Tao <tao.xavier@outlook.com>

* fix(rbd): decode the SoA link workspace for the multibody readback and drop the stale set_gravity copy

Completes #12

* perf(rbd): dedupe shared TriMesh uploads in from_rapier

Replaces #19

Co-Authored-By: Haixuan Xavier Tao <tao.xavier@outlook.com>

* perf(rbd): optional GPU contact reduction, merging per-pair manifolds to <=4 points

Replaces #17

Co-Authored-By: Haixuan Xavier Tao <tao.xavier@outlook.com>

* fix(rbd): pass the prediction distance to manifold_reduction in the contact-reduction kernel

Completes #17

* perf(rbd): flat 1-D narrow-phase dispatch, packing warps across batches

Replaces #21

Co-Authored-By: Haixuan Xavier Tao <tao.xavier@outlook.com>

* fix(rbd): restore the contacts capacity binding and import atomic_load_u32 for the flat dispatch

Completes #21

* fix(rbd): drop the stale 2mm PREDICTION constant reintroduced by the flat-dispatch port

Completes #21

* fix mpm feature-gating

* feat(rbd): expose dof_state_mut, links_static, joint_constraints and link_of_body

* feat(rbd): env-reset primitives, GPU motor scatter, contact sensors, actuator delay, encoded step, substep-refresh cadence and per-DoF armature/frictionloss

* fix(rbd): guard against implicit-coriolis drifting from the batch_indices uniform

* feat(rbd): cluster contact manifolds by normal, matching rapier, with a tunable threshold

* feat(rbd): model multibody joint frictionloss as a constraint instead of a force

* feat(rbd): seed per-DoF joint friction from rapier's Multibody::frictions

* refactor(rbd): read the contact prediction distance from RbdSimParams instead of a dedicated uniform

* chore: cargo fmt

* refactor(rbd): move the contact merge cosine into RbdSimParams

* refactor: move read_multibody_links onto NexusState and drive every env from control_multibody_motors

* test(rbd): add a headless many-small-environments step-timing harness

* revert(rbd): drop the flat 1-D narrow-phase dispatch

Measured 7-21% slower on Metal.

* chore: cleanup comments

* fix: gate control_multibody_motors on dim3 so the 2D build still compiles

* fix instability in joint-ball3 demo

* chore: remove debug test files

* chore: cleanups

* fix(rbd): build the bench harness without the metal feature and only in 3D

* fix(rbd): silence the clippy needless-borrow and unnecessary-mut lints

* fix(rbd): split the joint-constraint back-solve into its own dispatch to fit 8 storage buffers

* test(rbd): keep the bench harness under wgpu's default buffer-size limit

* fix(rbd): split the batched env reset into pose and DoF passes to fit 8 storage buffers

* chore: switch to the published rapier version

* fix: make all envs share the same RbdSimParams

* chore: clippy fixes

---------

Co-authored-by: Haixuan Xavier Tao <tao.xavier@outlook.com>
Co-authored-by: Haixuan Xavier Tao <tao.xavier@1ms.ai>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants