ML+Coffee: Qwen Powered Robotics & Surgical Video to Op Notes
ML+X Networking & Coworking
Event Details
ML+Coffee offers a supportive environment to discuss ongoing ML/AI projects and share knowledge across campus. Whether you're looking for advice on applying ML/AI to your data, hoping to demo a tool or method, or interested in discussing a paper, ML+Coffee offers the perfect space. The majority of our attendees are applied practitioners from diverse fields, not AI/ML purists looking to critique. Coffee and pastries provided ☕, courtesy of our sponsors.
Where: Room 1145, Discovery Building (330 N. Orchard St.).
When: Monthly on Wednesdays, 9-11am CT. Fall 2026 dates include: 9/9, 10/7, 11/11, 12/9.
Register: RSVP for ML+Coffee
Sept 9 Schedule
9:00-9:30 - Intros & Resource Sharing
The first 30 minutes are typically focused on introductions and casual share-outs about new ML/AI tools or resources folks are using.
9:30-10:00 - OpScribe, Operative Notes from Surgical Video (Ruffin Bryant)
A demo of OpScribe, which turns intraoperative video into a draft operative note and supporting CPT codes. De-identified video goes in, the system segments it into surgical phases, pulls tools and anatomy per phase, and writes a structured note where every claim ties back to a timestamped frame. Ruffin will run one case end to end, showing the surgeon-facing side (upload a case, watch the phase timeline populate, click any line in the note to jump to the frame supporting it) and the billing view, where candidate codes carry the same provenance link so a coder can verify rather than trust. He'll also cover the study underway with UW Health and how OpScribe benchmarks against general-purpose VLMs on tool, anatomy, and phase recognition.
10:00-10:30 - Driving a Robot Car with a Small Qwen Model (Runxiong "Calvin" Wu)
Meet Peter, a compact robotic car powered by an open-weight, 9B-parameter Qwen model. Served through vLLM on a lab workstation with an NVIDIA RTX 5090, the model translates camera images and spoken or typed instructions into short driving commands, which are sent to the car over Wi-Fi. Peter combines a CSI camera, a Qualcomm RUBIK Pi 3, and a WAVE ROVER chassis, with a total hardware cost of roughly $410-$480, excluding the GPU workstation. Calvin will demonstrate the car's "look-move-look again" feedback loop and show how a human-provided in-context correction, specifying a 45-degree left turn, let the car complete a harder task - locating a water bottle positioned behind it. Learn more at the QualcommClaw project site.
10:30-11:00 - Open Slot (TBD)
Have a demo, paper, or ongoing project to discuss? Add your name to the discussion queue. No formal presentation is required - this event prioritizes open dialogue and casual discussion. If helpful, bring a couple of slides or visuals (e.g., to share data, methods, or results). Many participants just bring a few key points or a rough overview of their work, and the conversation flows naturally from there.
We value inclusion and access for all participants and are pleased to provide reasonable accommodations for this event. Please email endemann@wisc.edu to make a disability-related accommodation request. Reasonable effort will be made to support your request.