تفاصيل العمل
Implemented TD3 (Twin Delayed Deep Deterministic Policy Gradient) for autonomous obstacle avoidance and goal navigation on a real Yahboom robot in a ROS2 environment. Designed a 24-dimensional LIDAR state space with reward shaping for smooth, collision-free trajectories; achieved < 10ms inference latency during live deployment. Trained fully in Gazebo simulation and transferred policy to physical hardware — complete sim-to-real pipeline with no retraining.