By Natan Reddy, Principal and Charlie Kirn, Investor
AI across the industrial sector is reaching a tipping point – moving from pilots into production, appearing on factory floors, job sites, in heavy equipment, and across supply chain routing networks.
Across manufacturing, construction, transportation & logistics, and alternative energy, major enterprises are approaching and crossing the chasm from experimental pilots into active production environments. Against a backdrop of a growing global emphasis on enterprise AI spending (more than 70% of enterprise leaders are shifting budgets from exploration to production-grade deployments), 86% of industrial leaders are prioritizing AI deployment for near-term cost reduction and productivity improvement, according to Bain.
Nevertheless, while McKinsey recently reported that 62% of firms it surveyed are actively experimenting with AI agents, its report revealed that industrial categories (e.g. supply chain, logistics, manufacturing) had the lowest levels of AI adoption and scaling among all industries surveyed, pointing to the reality that generalized consumer models often aren’t cut out for industrial use cases.
These statistics point to a clear shift: Small models are a necessary path to deeper AI adoption in the physical economy.
2025 was a year that every industry ramped up on AI agents and aggressively benchmarked generalized frontier models. While consumer-friendly, generalist LLMs from the likes of OpenAI or Anthropic are useful for broad ideation, their unspecialized nature presents a hindrance to safety, accuracy, and reliability across factory floors, construction jobsites, supply chains, and ruggedized environments.
Below, we explore and highlight what small models are, why they are a requirement for AI automation in mission-critical, hyper-specific industrial and physical environments, and key small model use cases across the industrial value chain (including a market map). Read on to learn more!
In this piece, we’ll use the term “small model” as an umbrella term for any light, compact, domain-tuned model built for specific, highly constrained tasks (e.g. technical manuals, SCADA logs, supply chain manifests, or CAD metadata, in the case of industrials).
Think of the difference between a foundation model and small model as an internal medicine doctor versus a specialized orthopedic surgeon. LLMs like GPT-4, Claude, or Grok, a subset of foundation models for example, is like a general practitioner – it knows a little bit about a lot. While helpful for brainstorming, summarizing broad text, or writing draft code, it comes alongside massive cloud compute, incurs high per-token costs, and reduced accuracy of output.
A small model like the examples listed above is the specialized orthopedic surgeon. While limited in broad general knowledge, it is built to excel at one thing: extreme, domain-specific precision. Instead of parameter counts in the hundreds of billions or trillions, small models typically range from 1 billion to 10 billion parameters. Because they are compact, they can run locally on an edge controller on a factory floor, on a ruggedized tablet on a construction site, or inside a micro-datacenter at a logistics hub.
Small models can include the following (please note some of these model types can also take on larger forms and are not exclusively ‘small’ e.g. sub 10 billion parameters).
Compact models focused purely on text and code. Instead of trying to know everything like giant cloud models, they are tuned for specific tasks—like reading technical manuals, writing reports, or summarizing system logs.
Models that serve as the brain for physical hardware and robotics. They take in visual feeds from cameras and plain-language instructions, then translate those inputs directly into physical movements like turning a valve, steering a vehicle, or picking up a part.
Lightweight visual processing models used to watch and analyze video feeds. On a factory floor or assembly line, they quickly identify items, track objects, and catch visual defects or damage in real time.
Models designed specifically to read continuous streams of numerical data over time. They monitor live sensor readings—such as temperature spikes, pressure changes, or electrical flow—to spot unusual patterns before equipment breaks down.
Fast-acting models programmed to make immediate operational choices. Rather than generating text, they analyze operational data and instantly trigger a specific command, like adjusting a thermostat, switching an API toggle, or applying an automated brake.
Ultra-tiny models compressed so small that they can live directly inside microchips, sensors, or basic circuit boards. They operate on minimal battery power to run simple, critical tasks—like shutting down a machine if dangerous vibrations are detected.
Don’t let the name fool you! The “small” in small model is just a nickname as a comparison to larger foundation models; it does not mean a small TAM. Small language models alone, a subset of our broader small model category defined above, are expected to grow to a global market size of $20B by 2030, according to Grand View Research. Looking through that lens alone is still not enough to get a sense of scale. If one looks at the global edge AI market (a broader representation of where small models are applicable), it is expected to grow from $30B in 2026 to a whopping $225B by 2035, according to Global Market Insights.
Today, vertical small models are already unlocking a massive addressable market in sectors that have historically resisted AI automation (for example, Siemens unveiled dedicated, task-specific industrial automation agents built specifically for the plant floor at 2026 Hannover Messe).
AI models can run the gamut from closed/proprietary to open weight to true open source. When it comes to small models in particular, it’s often imperative to be at least open weight to achieve full efficacy.
A few reasons for this, among others, include:
What small models lacks in broad applicability, it makes up for in number of other areas:
Small models win because their narrow scope enables greater reliability for deeply vertical tasks. This matches the level of accuracy required for mission-critical, safety-conscious tasks across the physical environment.
As small models are orders of magnitude smaller, running an inference query costs a fraction of a cent compared to routing that same query to a massive cloud LLM. When an industrial workflow requires thousands of automated decisions per hour, those unit economics make or break the ROI of an AI initiative.
Small models are small enough to run locally—directly on an edge controller, a ruggedized tablet, or an on-premise industrial gateway—with zero reliance on the cloud. This guarantees microsecond latency, total data privacy, and continuous uptime even when offline.
Having dozens of uncoordinated, generalist agents running around creates massive token bloat, duplicate data storage, and operational chaos. Enterprises are realizing they need a lightweight “Agent Fabric”—an orchestration layer that coordinates and governs specialized agents without burning through millions of tokens; small models fit cleanly into this strategy given their increased efficiencies.
Small models thrive in on-premise and local edge environments, allowing industrial companies to leverage their proprietary data without routing sensitive information through third-party cloud APIs. Combining open-weight small models with secure on-premise infrastructure (such as the Dell AI Factory with NVIDIA) and enterprise guardrail frameworks (like NVIDIA NeMo) gives organizations absolute data sovereignty, strict access control, and auditable compliance across sensitive operational environments.
Below we created a market map at the intersection of small models and industrial tech. To map the evolving landscape of specialized AI in physical industries, we categorized the ecosystem across the infrastructure stack required to train and deploy these models and the key operational verticals where they deliver immediate value. Please note this market map is non-exhaustive and categories are not mutually exclusive. Listed vendors were not independently verified as using small models; rather, they represent operational use cases uniquely suited for them.
The infrastructure layer shows how a model moves from training to deployment in industrial settings. Since model providers rarely offer off-the-shelf industrial solutions, fine-tuning, specialized hardware, and orchestration are essential to running AI across distributed physical environments. While some compact models, such as NVIDIA’s Nemotron series, directly address these constrained use cases, closed-weight providers like Anthropic and OpenAI are also featured given their broad market influence and their ability to adapt to industrial contexts through alternative training methods.
Manufacturing environments generate continuous streams of operational data, but the signal is rarely meaningful in isolation. A change in vibration, temperature, or pressure becomes actionable only relative to a particular machine, process, or operating baseline. This creates a need for vertically trained models that learn normal operating behavior and identify consequential deviation, rather than generalized models trained to recognize universal patterns. With unplanned downtime averaging $25,000 per hour, according to MaintainX’s 2024 State of Manufacturing Report, there is substantial operational value to unlock through the implementation of advanced models in this category.
Across factory floors, small models deliver immediate value in several core areas:
Construction environments combine changing physical conditions with project-specific terminology, document conventions, and operating workflows. Much of the underlying data is non-standardized, with inputs that vary by contractor, job type, and region. Models trained on these specific conventions and schemas outperform larger foundation models, which often misinterpret project-specific context or generate unreliable outputs.
Use cases range from document intelligence through autonomous control:
Efficient supply-chain operations depend on tightly coordinated physical assets across roads, warehouses, and distribution yards. AI workloads differ vastly by category: fleet platforms turn continuous telematics into operating decisions; warehouse systems make fast, repeatable physical-execution decisions; and autonomous trucks run a safety-critical perception-and-control stack onboard. In high-stakes environments where latency and efficiency dictate safety and costs, specialized edge models thrive by processing data locally, removing the heavy bandwidth demands of streaming video and telemetry to the cloud.
Small models solve these challenges in many ways:
Energy coordination increasingly requires specialized models as distributed generation turns a once-centralized grid into a network of assets with constantly shifting output, availability, and operating constraints. Effective models must be tuned to the physical and commercial realities of the system—context that general-purpose models do not inherently encode. According to Mckinsey’s June 2025 Report, U.S. power demand is expected to increase 3.5% annually through 2040, making precise coordination of distributed assets increasingly important.
Specialized models resolve these constraints across key energy operations:
In 2025, every industry jumped on the “agent high” and aggressively tested giant, general-purpose foundation models. But those generalist models are too slow and make too many mistakes for safety-critical jobs. While more than 60% of industrial companies are experimenting with AI, nearly two-thirds are stuck in the sandbox. They can’t scale because giant public models don’t understand specialized industrial language or manuals, and optimizing these systems to run within tight local hardware budgets is a massive engineering headache.
Letting dozens of uncoordinated AI agents run wild across a company creates operational chaos, redundant data, and massive bills. The huge opportunity here is building a lightweight “Agentic Fabric”—a centralized orchestration layer to manage, coordinate, and control all these specialized edge models safely and cost-effectively.
Closed-cloud models fail in the physical economy because they need constant internet, lag too much, and restrict customization. Open-weight and open source models are a necessity because they allow companies to keep their data completely on-site, giving them absolute data sovereignty.
We have to hammer this home again – small models do not equal small TAM or small impact. Small models are poised to capture market share from massively growing markets like edge AI, which is predicted to balloon from $30B to $225B over the next decade.
The real problem with this conversation isn’t the technology—it’s the marketing. Calling these incredibly powerful tools ‘small models’ does them a disservice. It makes them sound like a compromise, a diet version of real AI for companies that can’t afford a huge bill. But you don’t call a surgeon’s scalpel a ‘small sword’ and you don’t call a jet engine a ‘small power plant.’ These aren’t small models; this is truly precision AI, where specialization, rather than parameter count, is the defining feature.