Engineering & Technologyarticle2026-09-02

Value-Gradient Generalist Design of Muscle–Tendon Parameters for Cross-Terrain Musculoskeletal Locomotion

Open access0 citations

Abstract

Compliant muscle–tendon mechanics can improve terrain adaptation in musculoskeletal robots, but heterogeneous terrains impose competing requirements on compliance, propulsion, foot clearance, and support transfer. This paper proposes Value-Gradient Generalist Design (VGGD), a framework for selecting one fixed physiological muscle–tendon parameterization for cross-terrain locomotion while preserving skeletal topology and muscle routing. VGGD searches a compact PCA-based manifold that coordinates bounded scale factors for muscle strength, contraction-velocity capacity, and passive elastic response. A design- and terrain-conditioned value proxy is learned from sampled latent designs and then optimized by proximity-regularized projected value-gradient ascent under a soft-worst objective. Independent proxy validation uses 50 random designs solely for calibration and 100 separately sampled test designs excluded from fitting and Stage-B optimization. On the 100-design test set, the proxy shows positive agreement with mean cross-terrain return (Pearson r=0.580, Spearman ρ=0.565) and worst-terrain return (r=0.583, ρ=0.551), with all bootstrap intervals above zero and Holm-adjusted permutation p=0.0006. A separate paired local-direction test uses 60 previously unused evaluation seeds and equal feasible-space perturbation radii; at the nominal, midpoint, and selected designs, the proxy-gradient direction agrees with improvements in mean and seed-wise worst-terrain rollout return. After design selection, the muscle–tendon parameters are fixed, and a terrain-aware controller is trained with variational information-bottleneck regularization and auxiliary expert distillation. Checkpoint-resolved evaluation records identify 216/300 successes and an overall mean distance of 12.14 m for the complete pipeline, compared with 57/300 and 6.20 m for nominal-body PPO. The three checkpoint success rates are 84%, 78%, and 54% for Proposed and 49%, 0%, and 8% for PPO; exact two-sided policy-level permutation tests yield p=0.10 for success rate and p=0.20 for mean distance. All methods receive the same nominal 100-million-step final-controller budget per training run, while the 10 M Stage-A budget and expert-pretraining costs are reported separately.

// Source

View paper (DOI)Open access versionOpenAlexBiomimeticsPublished 2026-09-02

Authors: Lidong Sun, Ye Wang, Hao Cha, Fuchun Sun

Institutions: Tsinghua University, Naval University of Engineering