Project
| # | Title | Team Members | TA | Documents | Sponsor |
|---|---|---|---|---|---|
| 19 | Soccer Tracking Laser Gimbal |
Aaron Zhang Colin Smithsuvan Jayden Kim |
Wesley Pang | ||
| \#\#\# Soccer Tracking Laser Gimbal Team Members: - Colin Smithsuvan ([colin9@illinois.edu](mailto:colin9@illinois.edu)) - Jayden Kim ([jaydenk2@illinois.edu](mailto:jaydenk2@illinois.edu)) - Aaron Zhang ([aaronz4@illinois.edu](mailto:aaronz4@illinois.edu)) \# Problem Typical sports broadcasting cameras are human controlled for tracking objects. We'd like to try to automate this with a vision-based system and track a soccer ball, specifically two people passing the ball to each other. This could help reduce human error and smooth out the motion of the camera and keep the object being tracked more centered in the filming camera's field of view. An extended scope of this project is tracking of an entire field. \# Solution We aim to solve this problem by having a camera positioned on a gimbal’s arm, which follows the movement of the soccer ball. This way, the ball is always positioned in the center of the camera frame, regardless of where the ball is being moved. For tracking the ball, we plan to deploy a vision detection model, which is trained to detect the soccer ball with real-time accuracy and latency. From there, we will have an arm with a camera at the base along with a laser which points at the ball for human verification. This arm will track the movement of the ball, which in turn will move the camera with it. This way, we can fully track where the ball is moving at all times with the ball in the center of the frame. \# Solution Components Laptop, Custom PCB, Motors, Power Supply \#\# Subsystem 1 Laptop - HP Pavilion Plus with RTX 4050 GPU and Intel Ultra 7 CPU with 1 TB storage and 32 GB of RAM. \#\# Subsystem 2 Custom PCB 1) Power 1) Based on motor nominal voltage, the input bus voltage/dc link voltage will be 24V 2) To power the CAN transceiver, there needs to be 5V as characteristic of CAN 3) To support the logic level components like the microcontroller, gate driver logic supply, ethernet PHY, and others there will need to be a 3.3V rail 1) There will potentially be both an analog and digital 3.3V rail where the two are separated by a ferrite bead to keep any switching or digital noise from propagating onto the analog rail 4) To generate the 5V, 3V3\_D, and 3V3\_A rails, we plan to not design a custom switching converter circuit and buy a power conversion “brick” 2) Microcontroller 1) STM32H753 1) Contains CAN/CANFD controller, ethernet MAC, 16-bit ADCs, 16 and 32-bit high-resolution timers 2) Depending on number of ADCs and GPIOs needed can select from the different pin packages 1) ADCs and GPIOs will be used for sensing circuits, encoder interfaces, or telemetry we want 3) Inverter 1) Gate Driver 1) Smart gate driver IC \- DRV8304 2) H-bridge FETs 1) N-channel fets for the three H-bridges \- will need to size the FETS to fit within the maximum phase currents of the selected BLDCs (1.4A) \- 06N06L 1) 60V VDS, 5.5A 4) Communications 1) CAN 1) TCAN3403DRQ1 2) Ethernet 1) PHY \- LAN8742 2) RJ45 Connector \+ Magnetics \#\# Subsystem 3 Motors - The motors will be used to drive the motion of the two axis gimbal and allow the object to remain near center of the camera’s FOV - Initially planned to use a brushless DC motor in order to learn the inverter stage design - Fallback plan to use an integrated servo or brushed DC motor if unable to get brushless design working - Motor will need an encoder interface for position tracking at low speed - Motor PN: DFRobot 4015 3-phase brushless motor - 24V, 1.4A \#\# Subsystem 4 Power Supply - Any power supply capable of 24V at \~10 A \#\# Subsystem 5 Firmware - State machine should look similar to this: INIT, CALIBRATE, IDLE, TRACKING, COAST (target lost, extrapolate briefly), SEARCH (scan pattern, FAULT) - Input capture on camera trigger latches encoder counts at exposure where the timebase is shared with the laptop - Potentiometer open-loop mode for inverter bring-up to test hardware before any vision/software code exists - Telemetry streaming at full loop rate - Drivers for Ethernet packing/unpacking, SPI, MAC, Auxiliary functions, timer, led controls, logic level gate control, etc \#\# Subsystem 6 Vision Model - Any on-device model for real-time streaming that can accurately provide bounding boxes. We will take these bounding boxes and take the middle of the bounding box for the location that the laser should point to. The SOTA model that we will try first is YOLOv26 since it provides fast on-device inference and is primarily trained on datasets such as COCO, which include a class for “sports ball.” - We may also need the camera intrinstics and extrinstics to map the image pixel coordinates to real world coordinates to understand how much to move the turret arm in the real world relative to the image detected bounding box. - In order to map from pixel coordinates to real world coordinates, our biggest challenge is accurate depth estimation which is required for perspective projection. Single camera depth estimation is an on-going research challenge with many models not being industry ready. Therefore, we may have to implement two gimbals for triangulation or to use a simple lidar for more accurate depth estimation for backup. However, our primary method of approach is to try monocular depth estimation models for our single camera use case. \# Criterion For Success Primary Objective: Demonstrate a closed loop, vision guided camera platform that keeps a moving target centered in frame. High Level Goals - Standalone vision model provides accurate bounding boxes on a soccer ball on laptop - Camera can successfully send gimbal position data to the MCU - MCU successfully ingests data and passes to the laptop - Laptop can run vision model inference and send position of where the gimbal should point back to the gimbal - Gimbal is able to move within a 120 degree field of vision both vertically and horizontally to match the position of the object detected - Keep a moving target centered in frame at all times autonomously - Gimbal is able to move fast enough to keep the target in frame. - Inference is developed with low-enough latency to have a fast closed-loop |
|||||