YouTube channel Modern Rogue ran an experiment that deserves more attention than it got. Hosts Brian Brushwood and Justin Young simulated Flock-style street camera angles, stripped the audio from their footage, and fed the silent video into publicly available lip-reading software, including a tool called Open-Alterego.
Under the right conditions, the AI read their lips. Not perfectly, and not from across the street, but well enough to reconstruct identifiable words and phrases from a conversation no microphone ever captured.
What the Test Actually Found
The results depended almost entirely on how clearly the camera could see the speaker’s face.
The conditions that produced usable results were specific: a large face in frame, frontal angle, and stable lighting. The subject also needed to be close enough to the camera that mouth movements registered with enough clarity for the model to work with.
Distance, side angles, and poor lighting all degraded accuracy sharply.
Where the tools recovered fragments of speech, post-processing made the output more useful. The team cropped footage to isolate a single speaker, ran multiple analysis passes, then fed partial results into a language model.
That model filled in the words the lip-reading AI flagged as ambiguous, inferring likely phrases from grammar and context. The conversation started coming back together.
“One issue that is often raised is the potential for malign surveillance, e.g. using CCTV to eavesdrop on private civilian conversations.” (Authors of “Sub-word Level Lip Reading With Visual Attention,” arXiv)
The Research Behind the Risk
Academic findings closely mirror what Modern Rogue observed in its informal experiment.
According to an analysis by sourcesecurity.com of automated lip-reading in controlled CCTV scenarios, systems can reach approximately 80% word accuracy for a single speaker when faces are clearly framed. That figure drops to around 60% with two speakers in frame.
Those numbers assume decent lighting, a frontal view, and a face large enough for the model to extract meaningful visual features from the mouth region.
Language models are doing significant work inside these systems. When a lip-reading model identifies some words confidently and flags others as uncertain, a language model steps in. It infers the missing pieces from probable sentence structure and common phrase patterns.
The result is a transcript that draws more from statistical prediction than pure visual evidence. It can still reconstruct the substance of a conversation.
The hardware is also advancing. Research using neuromorphic (event-based) cameras has demonstrated accuracy gains on lip-reading tasks in controlled settings. These sensors capture micro-movements of the mouth at higher temporal resolution than standard frame-based cameras.
Separately, research published in Nature Communications in 2022 showed that RF and Wi-Fi sensing can infer lip movements even through face masks. That extends speech inference beyond what any optical camera can see.
“The prospect of systems capable to read lips is attractive for a variety of real-life applications, like assistive devices for persons with disabilities, security and surveillance applications, or automatic transcription of video content.” (Bulzomi et al., CVPR Workshop 2023)
What Flock Cameras Are Built For, and Where the Gap Is
Flock Safety’s published materials emphasize vehicle identification; the company does not market face-level or lip-reading capabilities.
Flock Safety markets its camera systems for license plate recognition and vehicle tracking. Fixed-focus cameras aimed at passing vehicles are unlikely to consistently produce the large, frontal, stable facial image that lip-reading models require.
The gap is real, but the trend line is not reassuring. PTZ cameras with strong optical zoom are now standard across a wide range of surveillance deployments, and resolution across the industry keeps climbing.
Where PTZ systems are deployed, a camera tracking a license plate can, under operator control, zoom into a face standing nearby.
Current laws governing surveillance largely address audio recording and facial recognition. Legal scholars and researchers have not yet clearly established whether silent footage that enables conversation recovery should be treated as eavesdropping under existing wiretap statutes. No identified regulatory framework applies specifically to visual speech reconstruction.
That gap exists now, while the technology is still limited. Closing it will be considerably harder once the technology is not.




























