Ruihao (William) Wu

All projects

· Internship

Finding highlights in long videos

A two-stage pipeline that decides which parts of a long video are worth sending to a large multimodal model.

Role
Solution Engineer Intern, Microsoft
Place
Shanghai
Period
June to August 2026
Tools
Python, ONNX Runtime, CLIP, Azure OpenAI

This project came from two months as a Solution Engineer Intern at Microsoft in Shanghai, from June to August 2026. The starting problem was cost: sending every minute of a long video to a large multimodal model is expensive.

The first stage runs locally on a CPU. It samples one frame per second, turns each frame into a compact image embedding with a quantized CLIP model, splits the video into segments, and ranks them. Only the segments it passes are sent to the large model.

The local ranking looked good on the 26 labels it had been tuned on and was much weaker on new footage. So the final version combines its score with the large model’s own triage score, instead of trusting either one alone.

The evaluation used blind candidate pools with source details hidden, human labels, and paired comparisons between models with confidence intervals.

In one comparison a different model ranked segments better, within noise, but kept fewer of the highlights people had picked and used far more input tokens, so I recommended keeping the current one.

Most of what I took from it is about evaluation: a score measured on the data a system was tuned on said little about how it would do on footage it had not seen.