Conversation

Jarkko Sakkinen

I made the most trivial video detection for ReadSeek just using Qwem3-VL-2B:

1. Chop fixed number of frames from the video. They are not in fixed time distance. Positions are adjusted by weighting them based on difference between the frames.
2. Scale them down to a lower resolution.
3. Make a new bitmap with key frames as tiles.
4. Run inference over the resulting bitmap.

Cheap and dirty but does the job for the scale and scope of the tool :-)
0
0
1