01
Chose Gemini 2.5 Flash as the AI engine, decided by its unique ability to access videos directly from a URL
PROBLEMOther LLMs need to fetch and process captions separately, and can't handle videos with no captions
Other models, even when handed a YouTube URL, can't reach the video itself, let alone its audio or footage. Gemini was the only model that accesses audio, video, and captions directly from a URL. Weighing the balance of cost, response speed, and accuracy, I settled on Gemini 2.5 Flash.
→Any video can be summarized, regardless of whether it has captions
02
Made summary generation on-demand rather than automatic on save, eliminating wasted API cost
PROBLEMNot every saved video actually needs a summary
I considered auto-generating on save, but decided not every video needs a summary. Instead, it generates in one tap at the moment the user actually wants it.
→Cuts unnecessary API calls to zero while still letting you read instantly the moment curiosity strikes
03
Structured tags in three layers — API, AI, and custom — putting searchability at the product's core
PROBLEMMany videos have no YouTube API tags at all, leaving search coverage full of gaps
Since API tags alone can only cover so much, I layered on AI-generated supplementary tags and the user's own custom tags for a three-tier structure. Tags are unified in English, and at search time the user's input is translated to English in the background before matching.
→A library where you can pull up any video without a miss, without thinking about language