Businesses are generating more video than ever, yet most of it remains unused. From decades of broadcast archives to store surveillance footage and production assets, vast amounts of video sit on servers without analysis. This growing pool of “dark data” is exactly what a new video data infrastructure startup aims to unlock.
InfiniMind, a Tokyo-founded company led by former Google executives Aza Kai (CEO) and Hiraku Yanagita (COO), is building infrastructure that converts massive volumes of unviewed video and audio into structured, queryable business data.
Former Googlers Saw the Video Data Infrastructure Gap Early
Kai and Yanagita spent nearly a decade working together at Google Japan, where they observed how rapidly video creation was outpacing companies’ ability to extract value from it.
“My co-founder and I saw this inflection point coming while we were still at Google,” Kai said. By 2024, advances in technology and clear market demand pushed them to build InfiniMind independently.
Kai explained that earlier tools could tag objects in frames but failed to understand narratives, causality, or answer complex questions across long-form video. For enterprises holding petabytes of footage, even basic insights remained inaccessible.
Advances in AI Enable Modern Video Data Infrastructure
What changed, according to Kai, was progress in vision-language models between 2021 and 2023. These advances allowed video AI to move beyond simple tagging toward deeper contextual understanding. Falling GPU costs and steady performance gains also played a role, but capability was the key barrier that finally fell.
InfiniMind recently raised $5.8 million in seed funding led by UTEC, with participation from CX2, Headline Asia, Chiba Dojo, and an AI researcher at a16z Scout. The company is relocating its headquarters to the U.S. while maintaining operations in Japan.
Early Products Show Commercial Traction
InfiniMind launched its first product, TV Pulse, in Japan in April 2025. The platform analyzes television content in real time, helping media and retail companies track product exposure, brand presence, customer sentiment, and PR impact. Following pilot programs, the company already counts wholesalers and media firms among its paying customers.
Its next product, DeepFrame, is a long-form video intelligence platform capable of processing up to 200 hours of footage to identify specific scenes, speakers, or events. A beta release is planned for March, with a full launch scheduled for April 2026.
Focused Enterprise Approach Differentiates the Platform
The video analysis market remains fragmented. While companies like TwelveLabs offer general-purpose video understanding APIs, InfiniMind is positioning its video data infrastructure specifically for enterprise use cases such as monitoring, safety, security, and deep video analysis.
“Our solution requires no code; clients bring their data, and our system processes it into actionable insights,” Kai said. The platform integrates audio, sound, and speech analysis alongside visuals and is designed to handle unlimited video length with an emphasis on cost efficiency.
Seed funding will support continued development of DeepFrame, infrastructure expansion, engineering hires, and customer growth across Japan and the U.S.

Don’t miss out on our latest news—follow us for the latest AI news, breakthroughs, and insights that matter.