Skip to content
botgigs

Launching soon. No card required.

[ blog / automation ]

Computer Vision Defect Detection: What a Pilot Really Costs

July 21, 2026 · 9 min read · by the Botgigs team

[ HIRE-BRIEF GENERATOR ]

hire
stack

brief.json

[ pre-generated sample ]

best-effort AI estimate, not a quote or a match

job

ticket_01

scope of work

who to hire

screen for

effort estimate

questions to ask your hire

Like the brief? Get matched to the right specialist when we launch.

A computer vision defect detection pilot on one production line typically costs $50,000 to $120,000 and runs 2 to 4 months, with a narrow single-defect deployment on hosted tooling achievable from around $15,000 to $40,000. Cameras, lighting and labeled images drive the budget far more than the model does. A pilot is worth running when you know your current human inspection escape rate, because without that baseline you cannot prove the system is better, and AI vision typically lands at 95 to 99 percent detection accuracy against roughly 70 to 80 percent for human inspectors under real production conditions. Last updated July 2026.

Manual visual inspection is one of the most expensive quality bottlenecks left on a modern line, and it is expensive in a way that does not appear on any single budget line. It shows up as escapes that become warranty claims, as inconsistency between shifts, and as inspectors who are accurate for the first two hours of a shift and less so for the last two. That is the gap a vision system closes, and it is why the payback periods reported for these projects tend to be short.

What the money actually goes to

The surprise for most first-time buyers is how little of the budget is model work. A representative single-line pilot breaks down roughly like this.

Image collection and labeling, 30 to 45 percent. You need enough examples of each defect class, including the rare ones, annotated consistently by someone who knows what counts as a defect. This is the line that blows up. If your defect rate is 0.5 percent, collecting a few hundred real examples of a specific defect can take weeks of production time on its own.

Cameras, optics and lighting, 20 to 30 percent. Including mounting, enclosures for the plant environment and the industrial PC or edge device that runs inference. Lighting is the single most underestimated item on the list and the most common reason a pilot underperforms.

Model development and evaluation, 15 to 25 percent. Training, tuning and, more importantly, building an honest held-out test set that reflects real line conditions rather than the images that were easy to collect.

Integration and operator interface, 10 to 20 percent. Connecting to the PLC or MES, deciding what physically happens when a unit is flagged, and giving the operator a screen that shows why. A detection nobody acts on is a detection that did not happen, so something has to route each flagged unit to the right owner with the evidence attached instead of leaving it in a log file.

The baseline question that decides everything

Before you spend anything, answer one question: what is your current escape rate, and how do you know?

Most plants cannot answer it precisely, which is understandable and also the whole problem. If you do not know how many defective units get past inspection today, you cannot say whether a system catching 97 percent of defects is a large improvement or a lateral move, and you cannot compute a return. Spending two weeks establishing that number, by double-inspecting a sample or reconciling against downstream returns, is the cheapest part of the project and the part that makes the rest defensible.

The second baseline is cost per escape: warranty, rework, scrap, and the customer relationship where relevant. Detection rate times cost per escape times volume is the entire ROI model, and it should be written down before anyone quotes you a system. The general method for fixing a baseline before the build rather than arguing about it afterward is set out in how to measure AI agent ROI.

Where pilots go wrong

Too many defect classes at once. A pilot that tries to catch eleven defect types needs eleven adequately populated datasets, and the rare classes will be starved. Pick the two or three that account for most of the cost. Adding classes later is straightforward; recovering a pilot that failed because it was spread thin is not.

Inconsistent labeling. If two inspectors disagree about whether a borderline scratch is a defect, the model learns the disagreement and its accuracy ceiling is set by that noise. Write the acceptance criteria down, have two people label an overlapping sample, and measure how often they agree before you train anything.

Changing conditions after the fact. A model trained under one lighting setup degrades when somebody moves a lamp, changes a conveyor speed, or switches to a supplier whose material has a different finish. Fix the imaging conditions during the pilot and treat any change as a re-validation event. This is textbook data drift rather than a broken model, and the distinction that tells you which one you are looking at is explained in model drift versus data drift.

Testing on the training data. Reported accuracy means nothing unless it comes from images the model never saw, collected on a different day. Insist on that, and insist on seeing the false negatives rather than only the headline number.

What a good pilot looks like

Eight to twelve weeks. One line, two or three defect classes, fixed lighting, an agreed labeling standard, and a held-out test set collected across different shifts. Run in shadow mode for the last few weeks, where the system makes calls and a human still decides, and compare agreement. The output is not a demo, it is two numbers you can take to a capital committee: detection rate against your baseline, and false positive rate, which determines how much extra rework the system will create if you act on it automatically.

Budget the cameras and the labeling honestly and the rest tends to follow. A pilot that produces those two numbers has done its job even if the answer is that the economics do not work on this particular line, because that answer costs $80,000 instead of the $400,000 a plant-wide rollout would have cost to learn the same thing.

Before you scope a pilot at all

Check that vision is the right tool. If the job is really reading text, codes or labels off a part, that is OCR, and it is cheaper and more mature, a distinction worth settling first in computer vision versus OCR. If a hosted vision API handles your defect class out of the box, test it for a few hundred dollars before commissioning anything custom.

Where custom is genuinely required, the full cost structure of production systems, including the annotation economics that dominate them, is broken down under computer vision development services. To scope the work and get matched to engineers who have deployed inspection systems on a real line, describe the defect and the line in the AI proof of concept hire brief.

[ Early access ]

Put this into practice.

Describe your automation in the free demo, get a scoped hire brief, and join early access to get matched at launch.

Launching soon. No card required.