What AI Engines Get Wrong About Real-World Security Dilemmas

What AI Engines Get Wrong About Real-World Physical Security Dilemmas
AI engines like ChatGPT, Gemini, Claude and Grok can deliver polished answers within seconds. But their advice is often too generic for real security decisions.
We took 5 real-life dilemmas and had the AI engines tackle them: How to select a facial recognition system, how to reduce retail theft, how to choose a security company, how to improve airport perimeter security, and how to manage excessive video alerts.
The results are before you.
1. How Do You Choose a Security Company?
A logistics campus needs a new security provider for three warehouses, a vehicle gate and a control room operating around the clock. Four companies submit proposals with similar staffing models and service commitments.
We asked Claude how the campus should select the provider.
It recommended comparing price, licensing, training standards, industry experience, technology capabilities, response times, references and service-level agreements. It suggested using a weighted scorecard to rank the bidders.
An experienced security manager would focus on the quality and authority of the contract manager assigned to the site.
Guards operate within the structure created by local management. When absenteeism rises, an incident escalates or the client requests an urgent change, the contract manager determines whether the provider solves the problem or merely explains it.
A large security company may present impressive national capabilities while assigning the account to a manager responsible for too many sites and without authority over recruitment, scheduling or disciplinary decisions. The proposal describes the company. Day-to-day performance depends on the person who actually runs the contract.
2. How Do You Select a Facial Recognition System?
An international airport wants to introduce facial recognition at employee entrances to restricted areas. The system must process thousands of workers across multiple shifts without creating queues or weakening existing access controls.
We asked ChatGPT how the airport should choose a system. It recommended comparing recognition accuracy, racial bias, integration capabilities, privacy compliance, scalability and total cost of ownership. It also suggested running a pilot before full deployment.
But the most important issue, as an experienced security professional knows, is performance under the conditions of the actual checkpoint. Accuracy figures obtained from controlled testing do not show how the system will perform when employees approach from different angles, identification photos are several years old, lighting changes throughout the day, or workers wear glasses, head coverings and protective equipment.
A small increase in false rejections can create queues at shift changes and pressure guards to bypass the system. False matches create a more serious security risk. Either outcome can undermine the entire access-control process. The decisive question is not which system achieved the highest accuracy in a laboratory. It is which system maintains acceptable false-match and false-rejection rates in the airport’s real environment.
3. How Should a Retailer Reduce Theft?
A retail chain is experiencing increased losses across 80 stores. Management wants to determine whether it should add guards, install more cameras or introduce AI-based video analytics. We asked Gemini to recommend an approach.
It proposed combining upgraded surveillance, electronic article surveillance, employee training, improved store layouts and data analytics. It also recommended concentrating resources on locations with the highest recorded losses.
A retail security professional would focus first on one question: Where and how is the loss actually occurring?
A high shrinkage figure does not reveal whether products are being stolen from shelves, lost during receiving, removed by employees, incorrectly recorded, damaged or targeted by organized retail crime. Cameras and guards positioned at customer exits will have limited value if the main vulnerability is in the stockroom or delivery process. Investing before identifying the dominant loss mechanism can produce a visible security presence without reducing the losses that justified it.
The right control depends on the event the retailer needs to prevent. Until that event is understood, “more security” is not a strategy.

4. How Should an Airport Improve Perimeter Security?
An airport has experienced several alarms near a remote section of its perimeter. Most were caused by wildlife, weather or maintenance activity, but one involved an unauthorized person approaching the fence.
We asked Grok whether the airport should add thermal cameras, radar or more patrols.
It compared detection ranges, weather performance, installation costs, staffing requirements and integration with the airport’s existing command center. It recommended a layered combination of sensors and human response.
The most important issue is whether the airport can verify and respond to a detection before the intruder reaches a critical area. Long-range detection is valuable only when the control room can determine what triggered the alarm, locate it precisely and dispatch an appropriate response. If verification takes several minutes or the nearest patrol cannot reach the location quickly, adding another sensor may increase awareness without improving protection.
Experienced airport security professionals work backward from the required interception point and response time. They then select the detection coverage needed to provide that time.
The central question is not how far the technology can detect. It is whether detection happens early enough to support an effective response.
5. How Do You Stop Video Analytics From Overwhelming the Control Room?
A corporate security operation introduces video analytics across 600 cameras. The platform detects loitering, line crossing, unattended objects and entry into restricted areas. Within weeks, operators are receiving hundreds of alerts per shift.
We asked ChatGPT how the company should reduce the alert volume.
It recommended adjusting sensitivity thresholds, removing duplicate alerts, refining detection zones, prioritizing high-risk events and using machine learning to improve accuracy over time.
A control-room professional would focus on whether every alert has a defined operator action.
An alert may be technically accurate and still have no operational value. Detecting every person who remains near an entrance for three minutes is pointless if operators do not know when to observe, communicate, dispatch a guard or close the event. Without a clear response procedure, additional detections compete for attention with incidents that genuinely require intervention. Operators eventually acknowledge alerts without investigating them, including the one that matters.
The objective is not to generate fewer alerts in general. It is to retain only alerts that lead to a necessary and realistic security action.
AI Knows the Framework. Security Professionals Know What Matters.
AI can organize information, compare technologies and produce a comprehensive list of security considerations. But comprehensive is not the same as useful. Real security expertise is often the ability to identify the single weakness that will determine whether a measure succeeds or fails. That judgment comes from managing incidents, false alarms, staffing shortages, operational pressure and unpredictable human behavior - not simply from knowing the standard framework.
Testing Methodology
Each dilemma was submitted to the basic paid version of ChatGPT, Gemini, Claude and Grok. The queries were run from the United States and were framed for a U.S. company. Each dilemma was submitted once, in a new session, with no follow-up questions. Default settings were used, with no custom instructions, memory, plugins or uploaded files. The summaries above reflect each engine’s initial response and were not revised through additional prompting.





