· Internship
Finding highlights in long videos
A two-stage pipeline that decides which parts of a long video are worth sending to a large multimodal model.
- Role
- Solution Engineer Intern, Microsoft
- Place
- Shanghai
- Period
- June to August 2026
- Tools
- Python, ONNX Runtime, CLIP, Azure OpenAI
This project came from two months as a Solution Engineer Intern at Microsoft in Shanghai, from June to August 2026. The starting problem was cost: sending every minute of a long video to a large multimodal model is expensive.
The first stage runs locally on a CPU. It samples one frame per second, turns each frame into a compact image embedding with a quantized CLIP model, splits the video into segments, and ranks them. Only the segments it passes are sent to the large model.
The local ranking looked good on the 26 labels it had been tuned on and was much weaker on new footage. So the final version combines its score with the large model’s own triage score, instead of trusting either one alone.
The evaluation used blind candidate pools with source details hidden, human labels, and paired comparisons between models with confidence intervals.
In one comparison a different model ranked segments better, within noise, but kept fewer of the highlights people had picked and used far more input tokens, so I recommended keeping the current one.
Most of what I took from it is about evaluation: a score measured on the data a system was tuned on said little about how it would do on footage it had not seen.