FieldVox: the offline voice assistant for field technicians
The problem
Factories, mines, oil fields, and rural clinics have no reliable connectivity, which structurally excludes every cloud voice assistant and copilot from the places technicians need answers most. The skilled-technician shortage (2.1M US manufacturing jobs projected unfilled by 2030) forces junior techs into senior-level diagnostics with nothing but a PDF manual. Existing industrial voice tools are scripted command-and-response systems that give orders, never answer questions.
The idea
Why now
The components for a fully offline spoken assistant only just fit on cheap hardware together: microcontroller-class ASR that beats Whisper-tiny, a 0.8B decision model in a 775MB CPU-friendly package, and on-device omni-modal models built for vehicles and edge. Builders are already prototyping fully-local voice-plus-LLM pipelines with thousand-comment threads, and the rugged-device market they would ride on is large and growing. The alternative incumbents all require cloud or six-figure rollouts, leaving the offline slot empty.
What it combines
Combines oido (Whisper-tiny-beating ASR on a $5 microcontroller) + gutsy-0-8b-open-decision-model-that-runs-on-cpu (0.8B calibrated decision model, 775MB GGUF, CPU-only) + banma-autoomni-2-0-23b-a3b-on-device-omni-modal-model-for-cars (on-device omni-modal model for edge hardware) + qwen-audio-3-1-realtime-plus-full-duplex-voice-model-with-think-act-speak-loop (full-duplex spoken interaction loop). Each alone is a component; together they make a fully offline spoken troubleshooting assistant viable on cheap hardware for the first time — the product is the integration, not any single model.
MVP
Build: a Raspberry Pi 4 with USB mic/speaker and a push-to-talk button, running an open tiny ASR model plus a small on-device LLM, with one 50-page equipment manual converted to a local retrieval index. Ask a repair question by voice, get a spoken answer grounded in the manual. Deliberately skip: wake-word detection, full-duplex barge-in, TTS naturalness, multi-manual support, any cloud component, and hardware miniaturization — the Pi proves the loop, not the form factor.
Distribution
The buyer is not the technician — it is whoever already sells into the site and owns the support-cost problem. Primary: industrial equipment OEMs (pump, compressor, genset, HVAC manufacturers) bundling the assistant as a pre-loaded diagnostic copilot on service tablets, turning their own manuals into answered questions. Secondary: rugged-device makers (Zebra, Honeywell, Getac) needing software differentiation on commodity Android hardware, and large facility operators or mining/oilfield services buying per-site licenses. Distribution piggybacks on hardware channels that already have the install base.
Why it wins
Picovoice sells licensed toolkit primitives, not a ready diagnostic product, and requires internet for licensing; Honeywell Vocollect is scripted command-and-response tied to expensive WMS rollouts — it gives orders, never answers questions; Augury is a cloud sensor-monitoring service with no voice interface and no dead-zone operation; Siemens Industrial Copilot is a cloud, Azure-backed enterprise play. Nobody ships an offline, voice-first, generative troubleshooting assistant.
Risks
Answer reliability is the existential risk: a confident-sounding hallucination in a safety-critical repair context destroys trust and creates liability. The MVP de-risks it by constraining the LLM to retrieval-only answers from the single manual, with a mandatory 'not found in the manual' refusal path — the 50-question accuracy/refusal evaluation is the actual deliverable of the weekend.
Build it with
- Oído — Whisper-tiny-beating ASR on a $5 microcontrollerWhisper-tiny-beating ASR on microcontroller-class hardware for the offline speech input layer.
- Gutsy: 0.8B open decision model that runs on CPU0.8B calibrated decision model running on CPU for answer/refusal judgments and intent classification without the cloud.
- Banma AutoOmni 2.0-23B-A3B: on-device omni-modal model for carsOn-device omni-modal reference for the fully-local speech-in/speech-out loop on edge hardware.
- Qwen-Audio-3.1-Realtime-Plus: full-duplex voice model with think-act-speak loopFull-duplex think-act-speak loop as the interaction model for natural spoken troubleshooting.
Repo to start from
fieldvox-offline — an open reference implementation of the MVP: Pi image build scripts, the ASR plus LLM pipeline, manual-to-index tooling, and the 50-question accuracy/refusal evaluation harness.
Evidence
- NAM: 2.1 million manufacturing jobs could go unfilled by 2030
- Offshore Technology: field connectivity gaps cost wrench time in oil and gas
- DataHorizon: rugged mobile hardware market sizing
- VentureBeat: Sonos acquires on-device voice startup Snips for $37.5M
- Siemens: Industrial Copilot wins Hermes Award 2025
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.