Computer vision defect detection does not output a binary pass/fail signal — it outputs a confidence score between 0 and 1 that tells you how certain the model is that what it detected is a real defect of a specific type. That score is the most important number in the AI-to-CMMS pipeline, because it determines whether a work order fires, what severity level it carries, and how much a technician should trust the detection before spending time and parts on a repair. This technical guide to computer vision defect confidence scores in CMMS explains what confidence scores measure, how to set work order thresholds correctly for your defect types and asset criticality tiers, and how to use confidence score distributions to identify when your model is drifting. Teams that ignore confidence score tuning end up with either alarm fatigue from too many low-confidence fires or missed defects from thresholds set too high. Start a free OxMaint trial to configure confidence thresholds for your AI inspection feed, or book a demo to walk through threshold calibration with your specific defect types.
Technical Guide · AI Vision · Defect Confidence Scores in CMMS
Computer Vision Defect Confidence Scores in CMMS
What confidence scores actually measure, how to set work order thresholds correctly, and how to use score distributions to detect model drift before it costs you missed defects or false alarms.
What a Confidence Score Actually Measures
A confidence score is the model's posterior probability that an image region contains a defect of a specific type, given all visual evidence in that region. It does not measure defect severity — it measures the model's certainty about the classification. These two dimensions require different threshold strategies.
0.0 – 0.50
Below action threshold — log as observation, no work order unless criticality is very high
0.50 – 0.75
Advisory threshold — trigger inspection work order; technician confirms or dismisses
0.75 – 0.90
Warning threshold — trigger planned corrective work order automatically
0.90 – 1.00
Alarm threshold — trigger immediate work order with emergency routing and parts check
Threshold Configuration by Defect Type and Asset Criticality
| Defect Type |
Critical Asset Threshold |
Standard Asset Threshold |
Low-Criticality Threshold |
| Structural crack |
0.60 → immediate WO |
0.75 → planned WO |
0.85 → inspection WO |
| Corrosion / surface degradation |
0.65 → planned WO |
0.75 → inspection WO |
0.88 → monitoring note |
| Seal / gasket leak |
0.55 → immediate WO |
0.70 → urgent WO |
0.80 → planned WO |
| Belt / conveyor wear |
0.70 → planned WO |
0.80 → inspection WO |
0.90 → monitoring note |
| Foreign object / contamination |
0.65 → immediate WO |
0.75 → urgent WO |
0.85 → planned WO |
Starting thresholds only. OxMaint tunes these automatically based on false positive / confirmed finding ratios from closed work orders over 60–90 days.
Configure confidence thresholds for your specific defect types and asset criticality tiers.
OxMaint's threshold engine accepts confidence scores from any AI vision system and maps them to work order priority levels based on your asset hierarchy and criticality configuration — all without custom code.
Detecting Model Drift from Confidence Score Distributions
A healthy AI vision model produces a bimodal confidence distribution: most detections cluster near 0.90+ (clear defects) or below 0.40 (clear non-defects), with relatively few scores in the 0.50–0.70 ambiguous range. When the distribution shifts toward the middle, the model is drifting — a signal to retrain before false alarm rates climb.
Healthy Model Distribution
0.0 – 0.4
Non-defects (large cluster)
0.4 – 0.7
Ambiguous (small cluster — normal)
0.7 – 1.0
Confirmed defects (large cluster)
Drifting Model — Retrain Signal
0.0 – 0.4
Non-defects (shrinking)
0.4 – 0.7
Ambiguous (growing — alarm signal)
0.7 – 1.0
High-confidence defects (shrinking)
Expert Review
PD
Pradeep Deshpande
Computer Vision Engineer, Industrial Inspection Systems — 9 years
The confidence score is the most underused number in most AI inspection deployments. Teams set a single threshold at 0.80, fire work orders on everything above it, and then complain about false alarms six months later. The right approach is to differentiate thresholds by defect type and asset criticality — because a 0.65 crack score on a primary structural beam should absolutely fire an immediate work order, while a 0.65 surface discoloration score on a low-criticality utility enclosure absolutely should not. OxMaint's threshold configuration makes this differentiation possible without custom code, which is the main reason most teams can implement it in a week rather than a quarter.
Frequently Asked Questions
How do confidence scores differ between defect types, and does that affect how we set thresholds?
Different defect types have different inherent model confidence ranges. Structural cracks on steel surfaces typically produce high confidence scores because the visual pattern is distinctive. Surface corrosion in variable lighting conditions produces lower confidence scores even for severe defects. OxMaint calibrates thresholds per defect type to account for this variance.
Book a demo to see per-defect threshold configuration for your asset types.
Can OxMaint automatically adjust thresholds based on false positive rates from closed work orders?
Yes. When a technician closes a work order with a "no defect found" failure code, OxMaint logs it against the specific confidence score band that fired the work order. If false positive rates for a specific score band exceed a configurable threshold, OxMaint flags the threshold for upward review and can auto-adjust within configured bounds.
Start free to enable automatic threshold monitoring on your first AI inspection feed.
What should we do when the model regularly produces confidence scores in the 0.50–0.70 range?
A sustained cluster of scores in the ambiguous range indicates model drift — typically caused by changes in lighting, camera positioning, or equipment surface appearance that the model wasn't trained on. The immediate action is to route these detections to inspection work orders (not immediate repair) and collect technician confirmations to build the retraining dataset.
Book a demo to see the model drift monitoring dashboard.
How does OxMaint handle confidence scores from multiple AI cameras on the same asset?
OxMaint uses a score aggregation rule for multi-camera assets: the highest score from any camera for the same defect type within the time window is used as the trigger score, and all camera views are attached to the resulting work order. This prevents the same physical defect from generating multiple work orders while ensuring the highest-confidence evidence drives the priority assignment.
Start free to configure multi-camera aggregation for your asset layout.
Is there a standard confidence threshold that works across all defect types as a starting point?
OxMaint recommends starting with 0.75 as the planned work order threshold and 0.90 as the immediate work order threshold for standard-criticality assets with any defect type. Critical assets should start at 0.60 for immediate work orders. These are calibration starting points — the optimal thresholds for your specific equipment and environment will stabilise within 60 to 90 days of closed-loop feedback.
Book a demo to walk through the starting configuration for your asset types.
A confidence score without a threshold strategy is a number without a decision.
OxMaint maps every AI vision confidence score to the right work order type, priority level, and routing rule — configured per defect type and asset criticality, tuned automatically from feedback.