Opportunities in Edge AI infrastructure and applications

What once needed a rack of datacenter GPUs now runs on a fraction of the hardware. A quantized model withup to 31B parameters now handles reasoning workloads with constrained memory and outperforms GPT 3.5, a 175B parameter model, at roughly one sixth of the parameter count. Capability at the edge is outpacing the deployment technology required in production. 

Small model development is outpacing the tooling to compress, ship, monitor, and update them. A team can build a model that fits on a device and still have no way to run ten thousand copies in the field, each on different hardware, each needing updates that sustain quality. Founders are starting those companies now, on the infrastructure side and the application side alike, and we want to meet them. These improvements are unlocking use cases that were impossible before, like robots that complete tasks successfully in new, unpredictable environments. To fully realize these capabilities at the edge requires deployment technologies to maintain production systems.

Capabilities are constrained by both hardware and algorithms

Today's frontier models are capable, but running them in the cloud is not always viable. A car has limited onboard compute and connectivity that drops offline, which makes depending on a large cloud model hard. Latency-sensitive tasks are vulnerable to the time that network round trips require, and workloads carrying sensitive data cannot leave the device at all. Cloud inference costs and data sovereignty concerns also push customers toward local inference, especially asmemory has gotten more expensive.

The mechanics are hard to fix with hardware alone. Generating a word means reading everything before it, so decoding is bound by memory bandwidth even though the first read, prefill, runs in parallel. A model that thinks through thousands of tokens cannot parallelize that thinking, and a device serving one person cannot batch like a datacenter does.Speculative decoding andprefix caching recover some latency, but they are runtime features, not the basis for a standalone company. Researchers we spoke with think these methods are close to commoditized.

Compression is an unsolved problem again

The remaining opportunity has less to do with standard quantization and more to do with making these techniques reliable, automated, hardware-aware, and accessible across models and devices. Naive compression can do outsized damage to reasoning models. But quantization-aware training and self-distillation now hold onto most of the original performances.

Research is converging on three possible solutions. The firstprotects the small share of weightscarrying the reasoning signal,as little as 2% in some studies, rather than quantizing everything at one rate. The secondcompresses the reasoning trace instead of the model,cutting cache size in exchange for some accuracy. The third patches specific failure modes, includingquantization-induced overthinking, where a model reaches the right answer mid-chain and then talks itself out of it. The strongest small open weight models, such as Gemma, nowapply targeted precision by layer, holding core reasoning at higher precision while compressing the layers that generate tokens.

A standalone compression company that lets a customer complete the same task at lower inference cost is very valuable today. Capable small models unlock edge workflows that cloud size and economics do not, which widens what a customer can automate and how quickly the model runs.

Three infrastructure areas to build in

  • Reasoning-preserving compression. We expect winners to package quantization-aware training and self-distillation into tools that work across models and hardware, so a customer gets a compressed model without needing a research team of their own. That is a valuable product today for customers that are already moving toward edge computing.

  • Fleet control and observability. Running thousands of deployed devices is a distinct problem from building models that run on them. Versioning, over-the-air updates with rollback, telemetry, and routing between on-device and cloud inference all have to work as a single system. This aligns with the cloud era, when monitoring servers and databases became a category of its own. Over-the-air delivery and device configuration follow conventions a competitor can absorb. Model versioning, quality regression detection, and routing are more AI native gaps that require modern solutions. AI is expanding which devices are worth connecting as well. For example, a sensor that used to just log data can now run a model, which necessitates a deeper device management layer, and we expect this market to grow. 

  • Small models built for the device. Rather than compressing a frontier model, there are some companies focused on building small models from scratch. The competition includes small models from the frontier labs and these improve every few months, but integrating small models with customers directly as Liquid AI has done with Mercedes is a durable strategy. The teams that win this category will compete on benchmarks relative to size, but more importantly work with device makers to gain enterprise distribution directly. The core design question is what to strip out while keeping what a given use case needs. That differs by application, so we expect many small models to end up being installed, built on a generalizable core that is adapted at the edges rather than one massive model that fits every task. 

Opportunities for applications

Applications are a large opportunity that will require the scaling of the aforementioned infrastructure. Models are now capable enough to fix processes that used to be unreliable, or to do things that were not possible at all. As such, edge AI applications come in two forms. Retrofit markets already have computer onboard technologies, which gives a model provider an integration opportunity. And greenfield markets that have no incumbent system, letting  a new entrant own the device and the deployment layer together.

Retrofit. Manufacturing, logistics, and automotive companies are sensitive to latency, have to run offline, and already have compute onboard existing machines. Automotive is the leading vertical today, and it goes beyond in-car assistants into driver assistance systems, lane-keeping, and automatic braking. Camera-based defect detection took one major US steel producer from 60 to 70 percent human-inspection accuracy toover 98 percent, saving $2M a year in labor cost. The open question is whether model quality alone carries these markets or whether per-customer deployment nuance caps scaling velocity. Home devices are a consumer version of this problem where trust is a moat and inference stays on the device. Liquid AI has been a leader in the retrofit market, with their ongoing partnerships with OEMs which reinforces the opportunity here being small models that can embed in legacy systems.

Greenfield. Physical AI, meanwhile, represents a large greenfield opportunity, and teams that can compress large reasoning models and manage the deployment layer for a large number of units will benefit. Compression matters in the case of humanoids since they face unstructured environments with no advance training. The nearer term opportunity is embodied AI systems shipping today such as autonomous mobile robots and robotic arms, where the same compression, deployment, and net new model stack can get to market today. Avatar Robotics and Sereact are two companies that are vertically integrating a hardware and edge device management approach with rapid deployment across warehouses, picking arms, and return processing stations.

Three Misconceptions

  • Edge AI is not always cheaper. Local processing cuts the recurring inference bill while the hardware costs money upfront. Companies with high volume workflows will benefit more from edge AI infrastructure as they will have a shorter payback period on the upfront hardware cost.

  • Edge devices are not solely reliant on local models. A connected device calls a cloud model when possible and most production systems route between the two.

  • Edge AI changes the information security landscape rather than removing it. Cloud data faces remote attacks while a seized device can be opened and data can be read off.

Get in touch!

We want to meet founders working on reasoning-preserving compression, fleet infrastructure for models already in the field, and models designed for the device from the first line of code. If that is you, reach out

Jason Rubenstein

Jason Rubenstein is an MIT Sloan MBA Candidate and a Summer Associate at Flybridge.

Next
Next

An Operating System for the ADHD Brain, and a Home for Its Rumination