Top 10 Mistakes AI Makes When Analyzing Interviews
At Satrix, we believe the right model for B2B interview research is not AI-only and it is not human-only. It is human-led, AI-assisted, and designed so that technology accelerates insight without pretending it can replace judgment. We are not skeptics of AI. We use it daily, and we have seen what it can do well; and how clients too quickly assume a short query in a chat bot can lead them to make overreaching strategic decision about their organization. This piece is about understanding the limits of the technology as it stands today and using it in a way that does not quietly degrade the value of the work, especially when the deliverable is being used to make consequential decisions about sales strategy, retention, product, or customer experience.
The most useful way we have learned to think about it is this: AI does not reduce the amount of human work involved in good interview analysis. It changes where the human work happens. The first couple of passes of scoping and pattern-finding are genuinely faster. But as the technology gets more capable, it creates new categories of complication that still require humans to manage. The QC work moves. It does not disappear. Treating that shift as a cost reduction is where most B2B AI rollouts go wrong.
There is also a more practical issue that the AI-moderated interview pitch tends to ignore. Put yourself in the reader’s seat and answer honestly: would your boss, or your boss’s boss, take a 30- to 45-minute call with an AI to talk through their buying decisions, their view of the competitive landscape, and their key strategies for the year ahead? If a 25-year-old researcher at a vendor’s firm cannot get that meeting on the calendar today, the AI version is not going to fare better. Senior executives give that time to peers who have done the homework and can hold the conversation at their level. That is not a problem AI can prompt-engineer 5-to-7 levels of questions its way out of in the next product cycle. It is a fact about who senior people are willing to be candid with.
The numbers around AI adoption support both sides of that view. McKinsey’s State of AI 2025 reports that 88 percent of organizations now regularly use AI in at least one business function, and 51 percent say they have already experienced at least one negative consequence from AI use, with inaccuracy the most commonly reported risk. Peer-reviewed research has also documented that automated speech recognition produces substantially higher word error rates for some speakers than for others, with averages of 0.35 for Black speakers versus 0.19 for white speakers across five major commercial systems (Koenecke et al., PNAS, 2020). Adoption is high, value is real, and so are the failure modes.
That distinction matters in B2B interview analysis, where a handful of conversations can influence sales strategy, retention plans, product priorities, customer experience improvements, and competitive positioning. When we run sales win-loss, churn, customer experience, and advisory board programs, clients are not paying us to summarize words. They are asking us to understand why those words were chosen, what tension sits behind them, and what the organization should do next. The ten mistakes below are the places where, in our experience, AI most often degrades the quality of that answer when it is left to work without a human in the lead.
AI Gets the Transcript Wrong Before Analysis Begins
The first mistake is painfully simple: AI often starts from a flawed transcript and then treats it as usable evidence. In real B2B interviews, people interrupt one another, use shorthand, trail off mid-sentence, reference internal projects by acronyms, and speak with different accents and audio quality. The PNAS work cited above showed materially different error rates across speakers, and the issue is not just an artifact of older systems. It is built into the conditions of real-world interviewing.
The complication compounds when more than one person is on the call. Conversational ASR performance degrades as overlap and speaker count increase, which means a transcript can be directionally wrong before the analysis layer even begins. If the assistant attaches a key complaint to the wrong stakeholder, misses a short but decisive comment from procurement, or mangles a competitor name, the downstream analysis can still look polished while being strategically unsound. The QC work has not gone away. It has shifted to verifying who said what, before any insight work begins.
AI is excellent at organizing language. It is still far less reliable at understanding why a senior buyer chose those words in that moment.
Evan Klein, Founder – Satrix Solutions
AI Treats a Flawed Transcript Like Objective Truth
The second mistake is subtler. Even when the transcript is imperfect, AI systems behave as though they were handed a clean, objective record. That false certainty is dangerous because interview analysis is only as strong as the evidence underneath it. A 2026 study from Cornell and Carnegie Mellon researchers (Kadoma, Shrivastava & Naaman, CHI 2026) found that error-prone subtitles consistently lowered viewer evaluations of both speakers and content. The takeaway translates directly to analysis: imperfect transcripts do not just lose words, they reshape how credible and coherent the speaker appears.
This is one reason we do not treat transcripts as self-validating evidence at Satrix. We treat them as a layer that must be reviewed, interpreted, and, when necessary, corrected before anyone draws strategic conclusions. In our interview work, the deliverable is not “whatever the software wrote down.” It is a validated understanding of what the respondent meant, how strongly they felt it, and why it matters to the business question at hand. A neat transcript is not the same thing as a true one.
AI Misses Emotional Nuance and Contextual Meaning
The third mistake is that AI often identifies what was discussed without understanding how it was experienced. A 2025 study published in Digital Medicine compared investigator-led analysis with LLM-assisted analysis of 33 cancer patient interviews and found a familiar pattern: the human-led approach surfaced psychosocial and emotional themes, while the LLMs excelled at structural, temporal, and logistical patterns but struggled with emotional nuance and contextual depth. The clinical setting is not B2B, but the asymmetry is the same one we see in our work.
In B2B interviews, dissatisfaction is frequently softened by professionalism. A buyer says implementation was “fine,” but the hesitation, the sequencing, and the surrounding detail tell you it was disruptive. A customer says support is “responsive,” but the broader context reveals they no longer trust the guidance. A human can quickly react to this and share understanding, where AI needs a beat before it reacts and then asks another question, taking the interviewee out of the emotion and timing of the situation, which prevents good follow-up and a deeper dive into the phrases used. If we flatten that into sentiment labels and theme counts, we risk missing the real story: whether the respondent was disappointed, resigned, politically cautious, or quietly signaling risk. Those differences change what leadership should do next.
AI Cannot Run the Interview the Way a Senior Executive Needs It Run
The fourth mistake follows from the executive-access problem flagged at the top of this piece. Even setting aside whether senior leaders will sit down with an AI in the first place, the deeper issue is that the most valuable B2B interview insight does not come from a tidy transcript. It comes from a skilled interviewer hearing the half-answer, noticing the pause, and probing in the right place at the right time. C-suite and senior business executives answer carefully, frame things diplomatically, and rarely volunteer the most important context unprompted. Drawing it out is its own discipline.
Recent research underlines how badly AI handles the moderator role specifically. A Harvard Business Review piece published in early 2026 compared more than 1,600 executives with 13 leading AI models and found that the AI systems used markedly different question mixes than human leaders, often overemphasizing interpretive analysis while underweighting productive and subjective questions. The result is conversations that look thorough but quietly steer outcomes and create blind spots. Salesforce’s State of the Connected Customer research lines up with the same picture from the buyer’s side: 86 percent of business buyers expect to be treated as trusted advisors rather than transactional contacts, and 80 percent of customers want human validation of AI outputs rather than autonomous AI decisions.
This is where humans are not a backstop. They are the front line. An experienced interviewer can read the room, recognize when an answer is incomplete, and adjust in real time. They can challenge gently, sit with silence, and pivot to a more productive thread when the original question is not landing. They can build the credibility a senior buyer needs in order to speak candidly about a competitor, a failed implementation, or an internal political dynamic. AI can summarize the text it receives. It cannot generate the text that should have been there, and it cannot earn the trust required to draw that text out of someone who has every reason to keep it diplomatic.
This is the clearest example of why AI does not reduce the human work in interview analysis. It just relocates it. Save time on coding and you still need that time, and more, in the interview itself if the conversation is going to be worth coding at all.
If the analysis cannot tell the difference between a polite answer and a dangerous answer, it is not ready to lead the conversation.
Evan Klein, Founder – Satrix Solutions
AI Smooths Over Disagreement Between Stakeholders
The fifth mistake is one of the most common in B2B: AI tends to over-summarize across stakeholders and produce a clean narrative where the truth is genuinely conflicted. Enterprise relationships are multi-threaded, and the disagreement between members of a buying group is often the insight. Gartner research released in 2025 found that 74 percent of B2B buyer teams demonstrate unhealthy conflict during the decision process, with buying groups now ranging from five to sixteen people across as many as four functions, each with different priorities and opinions. Buying groups that reach consensus are 2.5 times more likely to call the resulting deal high-quality.
When AI flattens those views into a unified summary, the analysis loses its most actionable signal. A churn risk that lives in the user community but not at the executive level is a very different problem than one that has reached the C-suite. A renewal that the champion is fighting for and the CFO is questioning is not the same as one where everyone is aligned. The work of preserving disagreement, weighting it correctly, and explaining what it means is still human work. AI does not save that time. It just moves it from synthesis to verification.
AI Compresses Themes So Aggressively That Outliers Disappear
The sixth mistake is premature compression. AI is extremely good at finding repeating language patterns. But leaders do not always need the most common theme first. Sometimes they need the weak signal that points to a looming product problem, a competitive threat, or a renewal-risk pattern that has not yet reached scale. A 2025 JMIR AI study comparing LLM and human thematic analysis identified this as a known failure mode of LLMs, describing how models can “focus too much on common or popular topics, missing out on niche or less frequently discussed topics” and how, without external knowledge of the domain, they tend to overgeneralize themes and miss specific subtopics or nuances.
This is one of the new categories of complication AI introduces. The faster the synthesis, the easier it is to skip the question of what deserves more weight than frequency alone would suggest. That judgment requires someone who knows the business, the market, and the strategic stakes well enough to recognize a quiet signal as important. The more capable AI gets at finding patterns, the more important it becomes to have a human deciding which patterns matter.
AI Invents Quotes, Evidence, and Support It Did Not Find
The seventh mistake should make every executive team cautious. A 2025 study in Scientific Reports evaluated an AI chat for thematic analysis and found that selected quotes were unreliable. The authors documented examples where phrases were changed, where two quotes were combined, and where wording was meaningfully altered, with overall quote-to-theme support sometimes falling below 50 percent.
That risk is not academic. In a B2B boardroom, quotes are often the most persuasive part of a readout because they make findings feel concrete and human. If a quote is fabricated, stitched together, or attached to the wrong theme, trust in the entire analysis erodes quickly. This is why we treat quotes as evidence, not decoration. AI can help us find candidate passages faster. It cannot be trusted to establish evidentiary integrity on its own. The polish of an invented quote makes verification more important, not less.
AI Sounds Certain Even When the Evidence Is Ambiguous
The eighth mistake is false confidence. AI systems are fluent by design, which means they often express tentative interpretations in language that sounds settled and executive-ready. That is dangerous in interview analysis, where ambiguity is normal. Sometimes a respondent is conflicted. Sometimes two stakeholders genuinely disagree. Sometimes a theme is emerging but not yet validated. The right answer in those moments is not more confidence. It is more transparency.
NIST’s AI Risk Management Framework frames trustworthy AI in terms that include valid and reliable, accountable and transparent, explainable and interpretable, privacy-enhanced, and fair with harmful bias managed. None of those properties happen automatically. They require human judgment about what to measure, what thresholds to set, and when an AI output should be flagged rather than presented. That is more of the work that AI is supposed to eliminate but actually relocates.
AI Misses B2B Context, Competitive Language, and Buying Dynamics
The ninth mistake is domain blindness. AI can parse language well while still missing what that language means in a specific market, category, or buying context. The same JMIR AI analysis cited above explicitly identified “lack of context” as a recurring LLM challenge, noting that without external knowledge or the ability to track long-term context, models “might misinterpret or miss certain topic nuances.” In B2B, that misinterpretation is rarely small. The same words can signal very different realities depending on the segment. “Integration” might refer to ease of deployment, architecture flexibility, partner ecosystem depth, or post-sale service burden. “Value” might mean price, speed to ROI, executive credibility, or reduced switching risk.
Without that business context, AI often treats adjacent concepts as interchangeable when they are not. Closing the gap is not a one-time prompt-engineering problem. It is ongoing translation work between what the model produces and what the market actually means, and that work scales with the complexity of the client’s business, not with the speed of the model.
AI should help us process more evidence. It should never become an excuse to listen less carefully.
Evan Klein, Founder – Satrix Solutions
AI Creates Privacy, Governance, and Actionability Risk
The tenth mistake is operational rather than interpretive, but it is just as important: organizations treat interview analysis as if it were only a summarization problem when it is also a governance problem. B2B interview transcripts often include commercially sensitive information, relationship history, pricing context, employee commentary, and account-specific risk signals.
NIST explicitly frames trustworthiness in terms that include accountability, transparency, privacy, and fairness, not just performance. That means the workflow matters as much as the model. Who can access the transcript? How are outputs reviewed? Can recommendations be traced back to source evidence? What happens when the model is wrong? At Satrix, we adhere to strict data protection policies aligned with GDPR, California privacy laws, and other standards, and we design deliverables so that clients receive transcripts, analyses, dashboards, and recommendations in secure formats. That governance discipline is not overhead. It is part of what makes insight usable.
The Honest Value Proposition
AI in interview analysis is closer to even-Stevens on value than the marketing suggests. The first couple of passes of scoping and pattern-finding are genuinely faster, and that is real. But the time saved there is largely spent on the new work AI creates: verifying transcripts, validating quotes, calibrating confidence, preserving disagreement, weighing weak signals, translating domain context, and managing governance. That balance will improve as the technology matures. It has not improved as much as the headlines imply.
The right model is not human-only and it is not AI-only. It is a partnership where the technology accelerates evidence handling and humans do what they have always done in this work: conduct the difficult interviews, read the room, weigh the ambiguity, and decide what the business should do next. AI changes what humans do in interview analysis. It does not change how much human judgment the work requires. And in the part of the work where senior executives are sharing what really happened on a deal, in a renewal, or inside an account, the human is not optional. The human is the reason the conversation happens at all.








