Learning-Based Control and Tactile Perception for Robotic Systems
| dc.contributor.advisor | Chen, Xu | |
| dc.contributor.author | Hu, Xiaohai | |
| dc.date.accessioned | 2026-09-16T18:31:34Z | |
| dc.date.issued | 2026-09-16 | |
| dc.date.submitted | 2026 | |
| dc.description | Thesis (Ph.D.)--University of Washington, 2026 | |
| dc.description.abstract | Robots deployed alongside people, whether humanoids, mobile vehicles, or manipulator arms, must close three feedback loops simultaneously. A perception loop reports what is happening in the surrounding scene and at the points of contact between the robot and the world. A planning loop converts task-level goals into trajectories that the robot's mechanism can actually execute. A control loop drives the robot to track those trajectories under disturbances, modeling errors, and embodiment constraints. The reliability of the overall system is bounded by the weakest of these three loops, yet the corresponding research communities have, by and large, developed their methods in isolation. This dissertation develops three contributions, each targeted at one of these loops, and unifies them under a recurring methodological theme: physics-grounded learning for robotic perception and control, in which the underlying physical structure of the task is exposed to the learning machinery rather than treated as a black-box input. The first contribution addresses the perception loop. It introduces a learning-based tactile slip-detection framework built on the Shannon entropy of the marker displacement field of a GelSight Mini optical tactile sensor. The choice of feature is deliberate: contact information is spatially inhomogeneous, and the entropy of the displacement field captures the physical signature of incipient slip rather than an object-specific texture cue. Across a dataset of ten everyday objects, the entropy feature set delivers 95.61% five-fold cross-validation accuracy and maintains above 86% leave-one-object-out accuracy on unseen objects -- a level of generalization that velocity-only features do not reach. The classifier runs within the per-frame budget of a 25 FPS sensor, enabling closed-loop grip-force adjustment in a manipulation experiment that retrieves a book from a shelf. A follow-on collaborative paper extends the binary signal to a slip-severity regression with continuous feedback control. The second contribution addresses the planning loop, taking as given that a controller can only be as good as the reference it is asked to track. It develops the TOL (Trajectory Optimization via L-BFGS) framework for an underactuated four-degree-of-freedom upper-limb rehabilitation exoskeleton. TOL casts the entire N-waypoint reference trajectory as a single 8N-dimensional nonlinear optimization and, rather than treating kinematics as a black box, embeds the optimization within the differentiable structure of the forward kinematics: it uses automatic differentiation in JAX for exact Jacobians near kinematic singularities and warm-starts from per-waypoint inverse kinematics. A pilot study with four healthy subjects shows a 4.6x improvement in end-effector tracking and a 33-70% reduction in motion jerk relative to Jacobian inverse kinematics and damped least-squares baselines. The third contribution addresses the control loop, where perception and planning must finally be realized on a full-body robot. It presents ARHumanoid, a deployment-oriented pipeline for semantically guided whole-body humanoid control from human motion. Rather than proposing a new standalone learning algorithm, this chapter studies how existing human-motion reconstruction, robot retargeting, reinforcement learning, and multimodal reasoning components can be integrated into a stable hardware system that respects the physical structure of the problem. The pipeline converts monocular human demonstration videos into deployable Unitree G1 behaviors through gravity-aligned motion recovery with GVHMR, morphology-aware retargeting with GMR, geometric contact correction for contact-rich reference motions, PPO-based whole-body tracking in Isaac Lab, and ONNX-based inference inside a 50 Hz policy loop supported by a 500 Hz ROS 2 servo loop. A central design choice is to separate slow semantic skill grounding from fast balance-critical motor execution: a cloud-hosted multimodal large language model maps a robot-mounted near-first-person RGB observation and a natural-language instruction to a discrete skill selection, while the selected motor policy executes locally using proprioceptive observations only. This separation keeps visual latency, cloud inference delay, and semantic uncertainty out of the high-frequency control loop. The chapter frames ARHumanoid as a carefully scoped deployment architecture, with explicit limitations: the current system uses a small skill library, assumes semi-structured object placements, and does not perform closed-loop visual servoing or online object pose estimation during manipulation. Although the three technical chapters study different robotic platforms, they are bound by one thesis: learning becomes more reliable when it is not treated as a purely data-driven black box, but is instead constrained by the physical and geometric structure of the robot-task interaction. In tactile perception, this means recognizing that contact information is spatially inhomogeneous and that different sensing regions contribute unequally to slip-related inference. In trajectory optimization, it means embedding learning and numerical optimization within the differentiable structure of forward kinematics and robot dynamics. In humanoid control, it means preserving the gravity-aligned and morphology-aware structure of human motion while separating slow semantic reasoning from fast balance-critical motor execution. Across all three, the central claim is the same: learning-based robotic methods become more reliable, more generalizable, and more deployable when they are grounded in the physics, geometry, and timing constraints of the platform they control. | |
| dc.embargo.lift | 2028-09-05T18:31:34Z | |
| dc.embargo.terms | Restrict to UW for 2 years -- then make Open Access | |
| dc.format.mimetype | application/pdf | |
| dc.identifier.other | Hu_washington_0250E_29499.pdf | |
| dc.identifier.uri | https://hdl.handle.net/1773/57831 | |
| dc.language.iso | en_US | |
| dc.rights | none | |
| dc.subject | Robotics | |
| dc.subject.other | Mechanical engineering | |
| dc.title | Learning-Based Control and Tactile Perception for Robotic Systems | |
| dc.type | Thesis |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- Hu_washington_0250E_29499.pdf
- Size:
- 15.79 MB
- Format:
- Adobe Portable Document Format
