Has the Pentagon given up on AI polygraph analysis of trustworthiness?
As of July, the Pentagon had been prepping artificial intelligence tools that analyze text, voices and faces to estimate trustworthiness, according to declassified presentation slides and prior Defense Department statements.
Now, Pentagon officials say that some of that AI is not part of the current revamp of the polygraph interview process, after Defense News asked for a response to concerns about the earlier statements and the slides.
On Wednesday, the Pentagon’s Defense Counterintelligence and Security Agency, or DCSA, which vets security clearance holders and applicants, said in an emailed statement that “active research and development” for the “Modernizing Polygraph” effort “does not incorporate generative super intelligence, also known as artificial intelligence,” that creates new content — such as emails, graphics or voiceovers.
Nor does the effort include “large language models,” meaning chatbots like ChatGPT and other programs trained on giant data sets to engage in human conversation, DCSA added.
And the research does not involve “facial action coding [of emotions], micro-expression analysis or vocal analyses,” the statement said. “While those technologies are not part of the effort,” DoD will “continue to monitor emerging scientific capabilities to evaluate their potential utility down the road.”
DCSA did not address other types of AI models that, according to the slides, researchers had been developing to scan spoken and written language for “deceptive speech” patterns.
Wednesday’s statement came after Defense News sought a response to criticisms from national security lawyers, scientists and AI auditors about AI’s ability to analyze deception — meaning not just to lie, but to do so knowingly and willfully.
They warned that AI-aided scanning of speech patterns, vocal stress or facial expressions to read minds lacks scientific backing, will compound errors and may violate civil liberties in a domain where few such rights exist.
Defense News chronicled their concerns and the Pentagon’s prior research and development of AI-aided lie-detection and emotion-estimation tools.
The recent DCSA statement, including the project name “Modernizing Polygraph,” contrasts with several earlier DoD statements, 2027 budget documents and DCSA slides that Defense News obtained through an open records request.
In July, a Pentagon official, who spoke on condition of anonymity to discuss the technology, explained the advantages of AI-driven sentiment and deception analysis.
“AI processes data in real time and enables standardized, objective data analysis,” the official told Defense News, via email. “AI-based analysis is a force multiplier for our human investigators, not a replacement.”
The official added, “The models are primarily trained on data from laboratory-controlled studies using volunteer participants.”
The DCSA slides, which include presentations from 2021 through 2025, show that the Air Force, DCSA and academic labs had been working on “sentiment analysis” AI models and “deceptive speech” AI-aided analysis since 2019, as part of a “Credibility Assessment Modernization” project.
In July, the DoD official described the AI models as part of a larger multi-year “transformation,” which the official referred to interchangeably as “Polygraph+” and Credibility Assessment Modernization.
Current budget materials price the Polygraph+/Polygraph Next project at $31 million, including costs for AI “scoring algorithms” and “decision aids,” thermal imaging tools to spot blood flow changes that can reflect stress and other non-contact sensors.

On Friday, DCSA declined to provide additional information, including the scope of the “Modernizing Polygraph” project and the decision not to include AI-based voice screening and facial expression analysis.
Trained on Twitter and Reddit
Over the summer, Farakh Zaman, a developmental leader at the Air Force Office of Scientific Research, elaborated on the appeal of AI-aided deception analysis.
“By processing subtle physiological, vocal and linguistic cues, these developmental algorithms may help establish a more objective baseline for investigators and pinpoint specific areas requiring further human-led clarification,” he said by email in July.
Wednesday’s DCSA statement said that the Air Force is not involved in the “active” research and development phase.
One undated declassified DCSA slide describes the Air Force amassing “deception-related” text to create a “deep learning model,” an algorithm that analyzes hordes of data to find patterns. The Air Force trained an early model on text from websites including Twitter, Reddit and other vocabulary data sets such as WordNet, according to the slide.
“This baseline training helps the algorithms understand natural, modern linguistic patterns, slang and everyday vocabulary,” Zaman explained in July. He acknowledged the need for “further rigorous testing, validation and a clear transition path” before fielding a prototype.
A separate DCSA slide, dated May 2025, contains a flowchart that depicts an avatar — a computer-generated image of an interviewer — in front of a security clearance applicant. Next to the applicant, a camera and microphone feed audio, video and speech-to-text data from the interview into various large language models. The recording devices and algorithms interact to gauge the applicant’s emotional state, while a human investigator observes and keys in questions for the avatar to ask.
DCSA’s Wednesday statement said the university that drafted the chart is not involved in the current research and development stage.
According to another undated DCSA slide, DoD has set a goal for natural language processing technology to interpret an interviewee’s feelings with 75% accuracy.
In July, the Pentagon official said by email that a human investigator remains essential for “the contextual reasoning and emotional intelligence required in these sensitive interviews.”
‘I’ve caught AI lying to me’
Ahead of Wednesday’s statement, Mark Zaid, a national security attorney who reviewed the slides, warned that the danger of derailing a military officer’s career is too great for AI to misinform an investigator.
The polygraph device “is just registering the physiology of the person,” such as spikes in breathing rate, perspiration or blood pressure, he said. “Then a human examiner analyzes those reactions to render an opinion.”
Here, however, “AI would be taking the polygraph readings and rendering an opinion” to the human examiner, said Zaid, who has represented individuals on all sides of the polygraph table, from workers disputing revoked clearances to former government officials and polygraphers.
“Now, we’re starting to get into the movie Minority Report,” he said, referencing a 2002 Steven Spielberg film that depicts a government reliant on psychic children for tips to predict crimes and detain people accused of future crimes.
Zaid, whose own clearance was revoked temporarily after he represented a whistleblower pivotal in President Trump’s first impeachment, said, “That concerns me more than the polygraph,” if science does not back AI’s reliability.
So far, scientific studies do not.
Rather, decades of research have discredited theories, popularized by shows such as Lie to Me, that anyone — human or artificial — can recognize deception based on speech patterns, vocal stress or facial movements.
Another concern scientists and auditors have is bias: the Justice Department has warned that using emotion-recognition algorithms to measure worker competency can discriminate against high performers with intellectual or developmental disabilities, who algorithms do not always understand.
Also, “automation bias,” the tendency to take AI’s advice without question even when conflicting information exists, may further erode the polygraph system’s reliability, technologists warned.
“[T]here is simply no consensus that polygraph evidence is reliable,” a 1998 Supreme Court ruling reads. Its decision banned polygraph results and polygraphers’ opinions from military courts.
Zaid said he appreciates the attempt to retool polygraph testing but wants proof that the planned approach is in fact an improvement.
“Anytime we might be able to scientifically advance the ability to determine truth, with accuracy, would generally be a good thing,” he said. But “the first thing that jumps out at me with AI is how often it is wrong…I’ve caught AI lying to me.”

Greg Rinckey, a former Army Judge Advocate General Corps officer who now provides security clearance representation, said that today’s “technology is so outdated, where they’re using tubes around people’s chests and blood pressure and sweat, so, any way that you can get a more reliable test for security clearances, especially counterintelligence investigations [into leaks], is a step up.”
That said, he questioned, “What is the research saying on how accurate it is?”
At present, research suggests the odds of AI-aided speech, voice or face analysis pegging a lie are about 50:50.
AI is ‘assuming dishonesty’
“There are no – reliable – diagnostic cues for deception detection, and I have yet to see any evidence, whether it is with humans or with artificial intelligence, that goes against that trend,” said David Markowitz, a Michigan State University professor who studies AI’s ability to understand human communication.
For instance, “the leakage idea — that these microfacial expressions will leak out from our awareness, signaling veracity or deceit — there’s very little evidence to support that, unfortunately,” he said.
Markowitz expressed doubts about AI-based deception analyses that claim an accuracy rate of 75% because “that is not how accurately humans do it, and often these systems are a big reflection of how humans do it.”
For example, he pointed to a recent experiment in which he prompted a Google Gemini AI model to examine audio and video of mock interrogations of accused cheaters in a trivia game, all of whom claimed innocence and some of whom were lying. Prior studies had asked humans to observe the same interrogations. DoD has entered into agreements to run a secure version of Gemini on unclassified and classified networks.
Markowitz said that his study, “The [in]efficacy of AI personas in deception detection experiments,” found Gemini was just as inaccurate as humans, both faring no better than chance, when prompted to distinguish fraudsters from truth-tellers.
What’s more, compared to humans, AI was more likely to wrongly accuse a trivia player of lying.
Markowitz gave one possible reason for AI’s systemic distrust of trivia players: The AI “is assuming that someone must have done something wrong because why else would they be questioned,” he said. “Anytime that someone is in an interview setting where they might be accused of something, a large language model is assuming dishonesty.”
Supreme Court: The jury is the lie detector
The consequences of a machine, whether a heart monitor or a voice screener, potentially bungling a credibility assessment can last a year or longer, said Zaid, a national security lawyer.
When officers who have successfully passed exams without issue suddenly fail their latest polygraph test, DoD typically puts them “in a corner with a dunce cap on for a year before officials will re-administer another exam,” he said. “You can’t get promoted. You can’t get certain assignments. And you can’t go overseas, which, to case officers, is death careerwise.”
Further, civil courts routinely dismiss complaints when employees claim that inaccurate results cost them their clearances. Nearly 40 years ago, the Supreme Court ruled that the courts have no power over such national security matters.
Concerns about power imbalances and due process prompted the European Union to label AI-aided polygraph tools as “high risk” and place strict limits on law enforcement’s use of them.
Stateside, the Supreme Court, in forbidding polygraph evidence in military courts, noted that allowing it would meddle with the human’s role, or, as it were, the jury’s role, in determining credibility.
“A fundamental premise of our criminal trial system is that the jury is the lie detector,” Justice Clarence Thomas wrote in the Court’s 1998 United States v. Scheffer decision.
Pointing to the court’s decision, Zaid said that he wants assurances of due process when staff undergo administrative proceedings to contest a revoked clearance resulting from misaligned algorithms. Comparing the situation to successful cases disputing speeding tickets caused by a miscalibrated street camera, Zaid questioned, “Would the person be provided the evidence? An understanding of the AI’s training?”
New biases
Pentagon officials have long stressed the desire for AI-based lie detection to help curb gender and cultural biases that human investigators may carry into an interview room.
At the same time, employers’ use of similar technology for pre-employment screening has faced allegations of other biases and system errors.
“Facial expression analysis is heaviest in terms of biases against individuals with disabilities,” said Jo Ann Oravec, a University of Wisconsin professor whose research focuses on AI, public policy and disability issues.
For instance, AI-powered video screening of job applicants tends to misinterpret the involuntary twitches or vocal outbursts of skilled people with certain disabilities — and thus pass over qualified applicants, said Oravec, author of the 2024 article “From Polygraphs to Truth Machines: Artificial Intelligence in Lie Detection.”
In response to concerns about such biases in July, the Pentagon official said via email that the credibility assessment modernization plan adheres to DoD’s AI Ethical Principles of “responsible, equitable and traceable” use of algorithms.
The Air Force’s Zaman, in his July email, said that his team’s work aligns with DoD’s Responsible Artificial Intelligence Guidelines.
Researchers “are actively engineering and testing these models” to account for “diverse, representative training data,” he stated. The testing covers “broad variations in speech, language and expression.”
Ryan Carrier, founder of ForHumanity, a global nonprofit that helps organizations audit AI systems, said that the Pentagon’s principles and guidelines are a skeleton of those that the European Union follows and that ForHumanity uses for inspections.
For example, the EU’s AI Act demands users of high-risk systems receive training to overcome a different kind of bias called “automation bias” — or the tendency to execute AI’s advice without questioning its accuracy. This bad habit can lead to incorrect data overriding correct decisions, Carrier said.
Also, multimodal systems, meaning AI systems that simultaneously parse audio (the subject’s voice), video (their facial movements) and text (interview transcriptions), tend to misalign data and results, yielding an irrelevant analysis, Carrier said.
The system “could be making determinations that a person is lying or trying to deceive me that are outside of the tool’s specifications,” he explained.
In response to concerns about multimodal systems in July, the Pentagon official had said, “Any automated output is treated strictly as auxiliary data for a highly trained, certified polygraph examiner.”
Similarly, Zaman, at the time, said that investigators who use AI-aided deception-detection systems are “trained to contextualize the tool’s findings, preventing automated outputs from dictating an outcome.”
To safeguard against mismatched data or results, the Air Force keeps a documented history of each piece of data’s origin and modifications and brings in independent white-hat hackers to try producing false correlations, he added.
“We recognize that multimodal AI systems present unique technical challenges, and mitigating these risks by hardening our systems is central to our engineering and testing protocols,” Zaman said.
Reason for suspicion?
In August, when the Pentagon administered a raft of polygraph exams to probe leaks alleging the Iran conflict is depleting U.S. arsenals, some observers suspected the endgame was not to find the leakers, but to crack down on disloyalty.
Critics of the planned system said that AI-based deception analysis may be best suited to achieve the same purpose.
“If it’s purely an intimidation technique, it may not be utterly useless from the point of view of large bureaucracies,” said Jay Stanley, a senior policy analyst with the American Civil Liberties Union’s Speech, Privacy and Technology Project.
The expectations for AI-powered trustworthiness testing resemble the hope that U.S. national security officials held out for the polygraph about 100 years ago, other lawyers said.
Polygraph devices “were seen as this magic-miracle-machine when they first came out,” and, over time, the government learned that when seemingly making deductions, the tools were “actually just flagging correlating factors, like heart rate and physiological factors that can indicate lying, but aren’t actually reflective of it,” said Jake Laperruque, deputy director of the Center for Democracy and Technology’s Security and Surveillance Project.
He added, “Thinking that AI sentiment analysis tools are actually making an assessment or engaged in reasoning — that’s not what those machines are actually able to do.”
Aliya Sternstein is an investigative journalist who covers technology, cognition and national security. She is also a research analyst at Georgetown Law. Her writing on the intersection of public health and constitutional rights has appeared in the Stanford Law Review, Arizona Law Review, and other law and health academic journals.