This project studies expressive robotic-arm motion in embodied AI using an FR3 robotic arm. The system generates physically feasible and human-interpretable motion in response to external signals, with music as the initial experimental condition. The framework includes three stages: human-motion retargeting for executable trajectories, cross-modal contrastive learning to align signal features with motion, and conditional generative modeling for coherent rhythmic action generation. Optional synchronized lighting may enhance presentation. Expected outcomes include a functional prototype, a complete signal-to-motion framework, and evaluations of executability, expressiveness, and interpretability for real-world human-robot interaction.