Founding AI Engineer (Computer Vision)
On-site · Engineering, Product & Tech · Full time San Francisco, California, United States
Our client is a stealth, seed-stage startup building an AI-powered wearable platform for industrial field technicians - running a real-time computer-vision and agentic AI system directly on smart-glasses hardware, live today with major data-center, energy, and industrial customers across multiple countries. Backed by a $5M seed round and founded by researchers from Harvard/NASA (computer vision, published at NeurIPS and ICCV) and MIT (machine learning, ex-quant researcher).
The bet: computer vision and agentic vision-language models (VLMs) grounded in real integrations with legacy enterprise systems (ServiceNow, SAP, Salesforce) - not another LLM wrapper. The industrial data this generates is also laying the groundwork for robotics automation down the line.
As Founding AI Engineer, you'll own the AI core: an agentic vision-language model (VLM) system doing multimodal visual reasoning, evals, and model orchestration, running on real hardware against real industrial workflows - not benchmark demos.
What you'll do
- Build and ship agentic VLM systems that reliably do visual reasoning in production (detection, segmentation, image/video understanding)
- Own model orchestration and build real evals discipline (ground-truth, trajectory, or regression harnesses)
- Work directly with the founding team, shaping the AI roadmap from day one
Location & culture
Fully on-site in the SF Bay Area. The founding team lives and works together for the first several months in a shared house before moving to a standard office - expect an intense, in-person, roughly 9-9-6 pace. Hours flex for exceptional, senior talent; the priority is getting the right person, not saving on salary.
Requirements
Essential
- 1–5 years applied AI/ML experience, ideally computer vision or multimodal — a practical builder, not a career academic
- Has shipped multimodal/CV systems to production in the VLM era, owning the model layer end-to-end
- Applied VLM/multimodal engineering specifically — vision-language, not sensor-fusion, audio-only, or time-series
- Applied agentic AI/model orchestration plus real evals discipline
- Startup or AI-team production experience, ideally as founder/CTO/founding engineer
Particularly valuable
- Edge/on-device inference; on-prem serving and fine-tuning (vLLM, SFT, quantization)
- RAG against knowledge bases
- Production AR/wearable AI or autonomous-driving CV experience
- Industrial domain exposure
Benefits
Compensation & benefits
- Base: $180K–$240K
- Equity: 0.25%–0.75%
- Visa sponsorship available (O-1, TN, E-3, H-1B1, H-1B transfer, F-1 OPT/STEM OPT)
How to apply: resume plus a short description of the most complex production VLM/CV system you've shipped.