sözaltı news Science
Science
EN AZ
AI turns bodycam footage into text. Can it keep the speakers straight?

AI turns bodycam footage into text. Can it keep the speakers straight?

phys.org 09.10.2026 15:00 8 views
Turning people into frogs—and back into princes—is the stuff of fairy tales. But back in January 2026, this fate befell a Heber City, Utah, police officer when AI accidentally cast the spell. Transcription software picke

This article has been reviewed according to Science X's editorial process and policies. Editors have highlighted the following attributes while ensuring the content's credibility: Turning people into frogs—and back into princes—is the stuff of fairy tales. But back in January 2026, this fate befell a Heber City, Utah, police officer when AI accidentally cast the spell.

Transcription software picked up the dialogue from "The Princess and the Frog" playing at the scene where the police encounter was unfolding. When AI software used the recording to draft the report, it mistook the movie dialogue for part of the encounter and claimed the officer had turned into a frog while on duty. The mix-up is a reminder that AI might take busywork off your hands, but a human still has to be at the helm.

This is exactly what Northeastern professor of criminology and criminal justice Eric Piza and graduate student Savannah Reid found when they set out to assess how well AI transcription software captured dialogue recorded by body cameras. As they report in their recent paper published in the Journal of Experimental Criminology, the bots whose work they reviewed got most of the text right. However, AI struggled to keep the speakers straight when they took turns.

The results show that human review remains essential before those transcripts are used to build a criminal case or hold an officer accountable, the researchers said. By pinpointing exactly where AI stumbles, the study helps investigators and researchers know what to check before relying on a transcript—for example, whether voices are missing or statements are attributed to the wrong speaker. The software identified fewer participants per transcript and assigned longer passages to a single speaker, losing some of the back-and-forth.

AI identified an average of 3.3 participants compared with 4.8 in the versions edited by people. It maxed out at seven, when its human counterparts identified as many as 16. AI transcripts were also watered down, containing about 25,000 fewer words across the sample.

Missing dialogue could mean losing an explanation, warning or response that helps establish why an encounter escalated. The researchers compared transcripts from 176 body camera videos that generated around 23 hours of footage covering 73 incidents. AI created one set, while human transcribers edited the other.

Extract — continue reading at the source.

Read full story