The Hardcore Moat of Machine Vision Lies in Digits Beyond the Decimal Point
I recently watched a video about machine vision, featuring a thought-provoking headline: "Principles Are Common Knowledge, Yet Precision Remains Hard to Match Nationwide."
To be honest, at first I thought it was just another cliché claiming "China cannot catch up." After watching through, I realized the story is far more nuanced.
When it comes to machine vision, everyone understands the basic principle — simply use a camera to capture images and run algorithms for recognition. OpenCV has been open-source for decades, and tutorials on deep learning are everywhere. It seems anyone can build a machine vision solution.
But once you step foot inside factories and onto production lines, you see a massive divide between building a functional system and building an excellent one.
Where does the gap lie? It boils down to the digits after the decimal point.
Let us first establish an intuitive frame of reference:
A human hair is roughly 0.1 millimeters, equivalent to 100 micrometers. Standard machine vision inspection delivers precision around 0.01 mm (10 μm), one-tenth the diameter of a hair. That already sounds impressive.
Nevertheless, high-end manufacturing demands far tighter tolerances — not merely 10 μm, nor even 1 μm:
What does 0.1 μm actually mean? If you slice a single human hair horizontally into 1,000 segments, each segment measures 0.1 μm.
At this precision tier, details routinely overlooked under normal conditions become fatal error sources:
The dividing line between seasoned machine vision engineers and amateurs is not whether defects can be identified, but whether precision can be steadily locked onto those decimal-place values across tens of thousands of inspection cycles.
An industrial camera is nothing like a smartphone camera.
Many people mistakenly believe algorithms form the core of machine vision. This is wrong.
The upper limit of any algorithm is defined by hardware performance. No algorithm, however sophisticated, can achieve submicron precision with a consumer-grade camera costing a few hundred yuan. It is comparable to trying to photograph cells under a microscope using a mobile phone camera — blurry images cannot be salvaged.
A complete industrial vision hardware stack consists of at least four layers: light sources, lenses, cameras and frame grabbers, each requiring rigorous optimization.
The most critical parameter for industrial cameras is not pixel count, but pixel size.
Consumer smartphone cameras chase high megapixel figures. A 100-megapixel sensor sounds compelling, yet individual pixels may be only 0.8 μm or smaller. Such pixels suffer low light intake, high noise and limited dynamic range. While adequate for portrait beautification, they are unsuitable for industrial inspection.
Industrial cameras follow the opposite design logic. High-end industrial sensors typically feature pixel sizes above 2.5 μm, and sometimes 5 μm or 10 μm. Larger pixels capture more photons, delivering cleaner images and more stable measurements.
Another critical specification is pixel depth. Ordinary consumer photographs adopt 8-bit encoding with 256 grayscale levels. Industrial inspection commonly uses 12-bit or 16-bit sensors, supporting 4,096 or even 65,536 distinct grayscale values. The difference may appear marginal at first glance, yet subpixel precision calculations rely entirely on this rich grayscale information.
If the camera defines the minimum achievable accuracy, the lens determines the maximum potential precision.
Images captured by ordinary lenses exhibit edge distortion: straight lines appear curved. Such distortion goes unnoticed by human eyes, yet becomes catastrophic for micron-level machine vision measurement.
For high-end industrial inspection, telecentric lenses are deployed instead of conventional optics. Within a defined working range, telecentric lenses maintain constant magnification regardless of object distance and achieve distortion rates down to several ten-thousandths or lower. As for pricing, a premium telecentric lens often costs more than the industrial camera paired with it.
At present, international brands still dominate the domestic market for high-precision industrial lenses.
A well-known industry maxim goes: "In machine vision, lighting accounts for 70% of success, algorithms only 30%."
Many inspection challenges stem not from flawed algorithms, but failure to properly illuminate defects for clear imaging.
Light angle, wavelength, uniformity and stability all shape final imaging quality. Engineers often spend far more time adjusting lighting hardware than tuning algorithm parameters.
Domestic manufacturers capable of stable mass production of high-end specialty light sources such as deep UV illumination remain scarce; most supplies still rely on imports.
Hardware establishes the baseline performance, while algorithms unlock the full potential of hardware. The pivotal technology here is subpixel precision.
What exactly is subpixel precision?
The fundamental unit of a digital image is a pixel. When the edge of a component crosses a pixel boundary, traditional algorithms can only confirm that the edge lies somewhere within that pixel, limited to integer-pixel resolution.
Subpixel algorithms leverage gradual grayscale variations and mathematical modeling to deduce the exact edge position inside a single pixel.
Achievable precision varies by algorithm:
What does 0.01 pixel mean in practice? Assume one pixel corresponds to a physical dimension of 10 μm. A precision of 0.01 pixel translates to a measurement error of 0.1 μm, or 100 nanometers. Without upgrading hardware, algorithms alone boost precision by a factor of 100.
Impressive as it sounds, this technology is hardly a secret. The underlying theory emerged decades ago, with abundant academic papers and open-source code publicly available.
So why does it remain a barrier? Precision quoted in research papers is measured under ideal laboratory conditions. Real-world production lines face constant interference:
Sustaining stable precision down to two decimal places under such chaotic conditions represents genuine engineering mastery. Consistency is ten thousand times harder than one-off precision.
This layer forms the true competitive moat.
Many outsiders simplify machine vision as "install a camera and write some algorithms." Nothing could be further from reality. Machine vision solutions built for different industries are essentially distinct technologies.
Additional application sectors include photovoltaics, automotive manufacturing, pharmaceuticals and food processing.
Optical schemes, algorithm architectures, mechanical equipment structures and validation standards differ drastically across verticals. Expertise gained in 3C electronics offers little advantage when entering semiconductor manufacturing. Mastery in these domains requires at least eight to ten years of industry accumulation.
Therefore, the core barrier in machine vision is not any single isolated technology, but integrated capabilities spanning optics, machinery, electronics, algorithms and vertical process know-how. No single link can be neglected.
Mid-to-low-end markets are largely conquered; high-end segments remain under active development.
To summarize: domestic suppliers have secured solid ground in mid-tier markets covering 3C electronics, photovoltaics and food packaging. Companies including Hikrobot, Dahua Technology, I-TEK OptoElectronics and TZTEK deliver competitive industrial cameras and vision systems at roughly one-third to half the price of imported alternatives, fully meeting standard inspection demands.
Founded by Dr. Dong Ning from the University of Science and Technology of China, I-TEK OptoElectronics developed China’s first domestically produced 8K line-scan industrial camera back in 2012, filling a critical domestic gap. The firm has since rolled out 16K line-scan cameras and 150-megapixel thermoelectrically cooled cameras, earning a seat at the global table for high-end industrial imaging hardware.
TZTEK achieves 0.3 μm inspection accuracy for consumer electronics, matching top-tier offerings from Hexagon and Keyence. Its wafer inspection solutions have reached nanometer-level precision, chipping away at KLA’s long-standing monopoly.
Nevertheless, substantial gaps persist within high-end fields, particularly front-end semiconductor inspection. Key bottlenecks include:
These overlapping constraints create the current market pattern: difficulty penetrating high-end sectors, while fierce price competition dominates mid-tier markets.
Returning to the video headline: "Principles Are Common Knowledge, Yet Precision Remains Hard to Match Nationwide."
This statement holds only half the truth. The theories — subpixel algorithms, telecentric optics, precision motion control — are all documented in textbooks. The real challenge lies in translating theory into stable industrial products, locking precision reliably to those decimal-place values, and bringing costs down to levels affordable for manufacturers. That is the true test of capability.
The machine vision industry has no room for revolutionary overnight breakthroughs or all-conquering "silver bullet" technologies.
Its competitive barriers are embedded within thousands of subtle details:
Such insights are rarely published in academic papers or fully disclosed within patents. They are forged year after year, line by production line, component by component.
China’s machine vision industry has developed over two decades: from complete reliance on imports, to full domestic substitution in mid-to-low-end segments, and now incremental breakthroughs at the high end. Progress has been arduous, yet the direction is clear.
After all, those tiny gaps after the decimal point can only be closed through persistent, incremental advances.
The Hardcore Moat of Machine Vision Lies in Digits Beyond the Decimal Point
I recently watched a video about machine vision, featuring a thought-provoking headline: "Principles Are Common Knowledge, Yet Precision Remains Hard to Match Nationwide."
To be honest, at first I thought it was just another cliché claiming "China cannot catch up." After watching through, I realized the story is far more nuanced.
When it comes to machine vision, everyone understands the basic principle — simply use a camera to capture images and run algorithms for recognition. OpenCV has been open-source for decades, and tutorials on deep learning are everywhere. It seems anyone can build a machine vision solution.
But once you step foot inside factories and onto production lines, you see a massive divide between building a functional system and building an excellent one.
Where does the gap lie? It boils down to the digits after the decimal point.
Let us first establish an intuitive frame of reference:
A human hair is roughly 0.1 millimeters, equivalent to 100 micrometers. Standard machine vision inspection delivers precision around 0.01 mm (10 μm), one-tenth the diameter of a hair. That already sounds impressive.
Nevertheless, high-end manufacturing demands far tighter tolerances — not merely 10 μm, nor even 1 μm:
What does 0.1 μm actually mean? If you slice a single human hair horizontally into 1,000 segments, each segment measures 0.1 μm.
At this precision tier, details routinely overlooked under normal conditions become fatal error sources:
The dividing line between seasoned machine vision engineers and amateurs is not whether defects can be identified, but whether precision can be steadily locked onto those decimal-place values across tens of thousands of inspection cycles.
An industrial camera is nothing like a smartphone camera.
Many people mistakenly believe algorithms form the core of machine vision. This is wrong.
The upper limit of any algorithm is defined by hardware performance. No algorithm, however sophisticated, can achieve submicron precision with a consumer-grade camera costing a few hundred yuan. It is comparable to trying to photograph cells under a microscope using a mobile phone camera — blurry images cannot be salvaged.
A complete industrial vision hardware stack consists of at least four layers: light sources, lenses, cameras and frame grabbers, each requiring rigorous optimization.
The most critical parameter for industrial cameras is not pixel count, but pixel size.
Consumer smartphone cameras chase high megapixel figures. A 100-megapixel sensor sounds compelling, yet individual pixels may be only 0.8 μm or smaller. Such pixels suffer low light intake, high noise and limited dynamic range. While adequate for portrait beautification, they are unsuitable for industrial inspection.
Industrial cameras follow the opposite design logic. High-end industrial sensors typically feature pixel sizes above 2.5 μm, and sometimes 5 μm or 10 μm. Larger pixels capture more photons, delivering cleaner images and more stable measurements.
Another critical specification is pixel depth. Ordinary consumer photographs adopt 8-bit encoding with 256 grayscale levels. Industrial inspection commonly uses 12-bit or 16-bit sensors, supporting 4,096 or even 65,536 distinct grayscale values. The difference may appear marginal at first glance, yet subpixel precision calculations rely entirely on this rich grayscale information.
If the camera defines the minimum achievable accuracy, the lens determines the maximum potential precision.
Images captured by ordinary lenses exhibit edge distortion: straight lines appear curved. Such distortion goes unnoticed by human eyes, yet becomes catastrophic for micron-level machine vision measurement.
For high-end industrial inspection, telecentric lenses are deployed instead of conventional optics. Within a defined working range, telecentric lenses maintain constant magnification regardless of object distance and achieve distortion rates down to several ten-thousandths or lower. As for pricing, a premium telecentric lens often costs more than the industrial camera paired with it.
At present, international brands still dominate the domestic market for high-precision industrial lenses.
A well-known industry maxim goes: "In machine vision, lighting accounts for 70% of success, algorithms only 30%."
Many inspection challenges stem not from flawed algorithms, but failure to properly illuminate defects for clear imaging.
Light angle, wavelength, uniformity and stability all shape final imaging quality. Engineers often spend far more time adjusting lighting hardware than tuning algorithm parameters.
Domestic manufacturers capable of stable mass production of high-end specialty light sources such as deep UV illumination remain scarce; most supplies still rely on imports.
Hardware establishes the baseline performance, while algorithms unlock the full potential of hardware. The pivotal technology here is subpixel precision.
What exactly is subpixel precision?
The fundamental unit of a digital image is a pixel. When the edge of a component crosses a pixel boundary, traditional algorithms can only confirm that the edge lies somewhere within that pixel, limited to integer-pixel resolution.
Subpixel algorithms leverage gradual grayscale variations and mathematical modeling to deduce the exact edge position inside a single pixel.
Achievable precision varies by algorithm:
What does 0.01 pixel mean in practice? Assume one pixel corresponds to a physical dimension of 10 μm. A precision of 0.01 pixel translates to a measurement error of 0.1 μm, or 100 nanometers. Without upgrading hardware, algorithms alone boost precision by a factor of 100.
Impressive as it sounds, this technology is hardly a secret. The underlying theory emerged decades ago, with abundant academic papers and open-source code publicly available.
So why does it remain a barrier? Precision quoted in research papers is measured under ideal laboratory conditions. Real-world production lines face constant interference:
Sustaining stable precision down to two decimal places under such chaotic conditions represents genuine engineering mastery. Consistency is ten thousand times harder than one-off precision.
This layer forms the true competitive moat.
Many outsiders simplify machine vision as "install a camera and write some algorithms." Nothing could be further from reality. Machine vision solutions built for different industries are essentially distinct technologies.
Additional application sectors include photovoltaics, automotive manufacturing, pharmaceuticals and food processing.
Optical schemes, algorithm architectures, mechanical equipment structures and validation standards differ drastically across verticals. Expertise gained in 3C electronics offers little advantage when entering semiconductor manufacturing. Mastery in these domains requires at least eight to ten years of industry accumulation.
Therefore, the core barrier in machine vision is not any single isolated technology, but integrated capabilities spanning optics, machinery, electronics, algorithms and vertical process know-how. No single link can be neglected.
Mid-to-low-end markets are largely conquered; high-end segments remain under active development.
To summarize: domestic suppliers have secured solid ground in mid-tier markets covering 3C electronics, photovoltaics and food packaging. Companies including Hikrobot, Dahua Technology, I-TEK OptoElectronics and TZTEK deliver competitive industrial cameras and vision systems at roughly one-third to half the price of imported alternatives, fully meeting standard inspection demands.
Founded by Dr. Dong Ning from the University of Science and Technology of China, I-TEK OptoElectronics developed China’s first domestically produced 8K line-scan industrial camera back in 2012, filling a critical domestic gap. The firm has since rolled out 16K line-scan cameras and 150-megapixel thermoelectrically cooled cameras, earning a seat at the global table for high-end industrial imaging hardware.
TZTEK achieves 0.3 μm inspection accuracy for consumer electronics, matching top-tier offerings from Hexagon and Keyence. Its wafer inspection solutions have reached nanometer-level precision, chipping away at KLA’s long-standing monopoly.
Nevertheless, substantial gaps persist within high-end fields, particularly front-end semiconductor inspection. Key bottlenecks include:
These overlapping constraints create the current market pattern: difficulty penetrating high-end sectors, while fierce price competition dominates mid-tier markets.
Returning to the video headline: "Principles Are Common Knowledge, Yet Precision Remains Hard to Match Nationwide."
This statement holds only half the truth. The theories — subpixel algorithms, telecentric optics, precision motion control — are all documented in textbooks. The real challenge lies in translating theory into stable industrial products, locking precision reliably to those decimal-place values, and bringing costs down to levels affordable for manufacturers. That is the true test of capability.
The machine vision industry has no room for revolutionary overnight breakthroughs or all-conquering "silver bullet" technologies.
Its competitive barriers are embedded within thousands of subtle details:
Such insights are rarely published in academic papers or fully disclosed within patents. They are forged year after year, line by production line, component by component.
China’s machine vision industry has developed over two decades: from complete reliance on imports, to full domestic substitution in mid-to-low-end segments, and now incremental breakthroughs at the high end. Progress has been arduous, yet the direction is clear.
After all, those tiny gaps after the decimal point can only be closed through persistent, incremental advances.