AI that holds in production.
I'm Oz Levi. Twenty years building data and ML systems. I get the data underneath right, so the RAG, search, and agent systems on top stay reliable, measurable, affordable, and ready to scale. Foundation first.
Most companies don't have an AI idea problem. They have an AI system problem.
The demo works. The prototype impresses the room. Then comes the hard part: making it scale, keeping it reliable, not blowing the budget. That gap between a promising proof-of-concept and a production-grade system is where most AI initiatives stall: RAND puts AI project failure above 80 percent, twice the rate of ordinary IT projects.
Closing that gap is my work, and I close it data-first: senior technical leadership and hands-on architecture, from the first honest assessment to a plan your team can build. Reliable AI is a system, not a model, and the data decides whether it holds.
Most consultants advise on the top layer. I work across all five.
The intelligence you see at the top is created at the bottom. Explore the layers to see where RAG actually lives.
Platforms, pipelines, streaming, lakehouse and warehouse design, data quality, lineage, governance, and privacy. The bedrock everything above depends on.
Classical and modern ML, model serving, MLOps and LLMOps, evaluation, observability, and cost control. The operational layer that turns a working model into a running product.
Knowledge graphs, embeddings, ontologies, semantic and dimensional modeling. How you structure and connect knowledge: the factual backbone for everything above.
Search, RAG, hybrid and semantic retrieval, GraphRAG, re-ranking, context assembly, and retrieval evaluation. High-performance retrieval is what makes AI factual.
LLM integration, agents, copilots, and decision engines. The visible product, only as reliable as the four layers beneath it.
“Just do RAG” touches layers 4–5. Reliability comes from layers 1–3: the work most teams skip. It's the work I've spent twenty years on.
Four ways to work with me.
Every engagement is scoped to where the real risk is, wherever in the stack that turns out to be.
Fractional CTO & advisory
Ongoing senior data/AI leadership: architecture calls, build-vs-buy, model and vendor selection, roadmap de-risking. Embedded as a fractional CTO or fractional AI engineer, on a retainer or a lighter advisory cadence.For companies that need a senior technical owner without the full-time hire.Retainer · ongoing02Architecture sprint
A specific system designed, prototyped, and de-risked: semantic search, RAG pipelines, AI agents, plus the data foundations, serving, and evaluation underneath. You get a spec your team can build from.For teams with a concrete system to design and the builders to build it.Fixed scope · weeks03LLM production-readiness audit
An honest expert review of a PoC or live LLM system before you scale it: reliability, cost, governance, evaluation coverage, observability, plus a prioritized roadmap to production-grade.For a system that works in the demo and now has to work for real.Time-boxed · roadmap04Technical due diligence
Independent assessment of an AI or data company for investors and acquirers. What's real, what's risk, and what it will cost to fix.For investors who need an engineer’s answer, not a pitch deck’s.Defined deliverablesemanticOps.ai is one person. Me.
When you hire semanticOps.ai, you get me: the architect who designs the system and stays through the build.
For 12+ years at Matrix DnA I was Chief Solutions Architect, then CTO, then EVP of Technology & Innovation, leading enterprise data and AI architecture and founding its Databricks Center of Excellence. I was running ML in production years before “GenAI” was a word.
Later I co-founded optmus.ai, building an operating system for AI. The point: I didn't arrive with the hype cycle. I arrived through twenty years of the data engineering that decides whether AI is reliable, or just a demo.
No bench, no juniors, no hand-offs. The person you talk to is the person who does the work: an engineer, not a slide deck. Based in Tel Aviv, working with companies worldwide.
Straight answers to the questions I hear most.
The PoC-to-production questions every team asks, answered the way I answer them on a call.
Why do most RAG systems fail in production?
Because the failure is almost never where teams look. When a RAG system hallucinates, the instinct is to tune prompts or swap models, layers 4 and 5 of the stack. In my experience the root cause usually sits lower: documents chunked with no regard for meaning, embeddings that flatten domain language, no evaluation harness to catch drift. The fix is representation and retrieval work. Model the knowledge properly, measure retrieval quality on your own data, then tune what the numbers say is broken.
RAG or fine-tuning: which one does my use case need?
Different tools for different problems. RAG injects knowledge at question time, so it wins when facts change often, when answers must cite sources, or when data is private and per-customer. Fine-tuning changes how a model behaves, so it wins for tone, format discipline, and narrow repeated tasks. Most production systems I design use retrieval for knowledge and reserve fine-tuning for behavior, if it appears at all. Start with retrieval: it is cheaper to iterate, easier to evaluate, and every answer keeps an audit trail of where it came from.
Is RAG dead now that models have long context windows?
No. Long context and retrieval solve different problems. A million-token window still cannot hold a corpus that changes daily, enforce per-user permissions, or tell you which document an answer came from. And stuffed context is priced per token, on every single call. Retrieval keeps the model grounded in a governed, current knowledge base at a fraction of the cost. What long context does change is chunking strategy: you can afford bigger, more coherent passages. That is an architecture decision, not the death of retrieval.
How do I improve RAG retrieval accuracy?
Measure first. Build a retrieval evaluation set from real user questions, score the system you have, and let the numbers pick the lever. The levers, in the order I usually reach for them: chunking that follows document structure instead of fixed sizes, hybrid retrieval that combines keyword and vector search, re-ranking on the candidate set, and richer representation, metadata and knowledge graphs, when entities and relationships matter. Most teams skip the measuring and tune blind. The measuring is the fix: each lever either earns its complexity on your data or it does not.
How do you evaluate an LLM application before launch?
With an evaluation pipeline built before the application is finished, not after. I start from a golden set: real questions with expert-approved answers, including the ugly edge cases. Retrieval gets measured on its own, does the right context surface, before generation is judged, is the answer faithful to that context. On top of that: automated scoring for regressions, human review where judgment matters, and observability in production so every answer can be traced back through the pipeline. If a system cannot be measured, it does not ship.
Why do most AI proofs of concept never reach production?
Because a demo only has to work once, and production has to work every time. The pattern repeats in every failed PoC I have audited: the data layer was improvised, nothing was evaluated against a quality bar, nobody priced tokens at scale, and there was no observability for when answers go wrong. Closing the gap means treating the PoC as evidence, not as version one. Re-architect on solid data foundations, build the measuring in from day one, then scale deliberately. Mapping that path is exactly what my LLM production-readiness audit delivers.
What does it cost to work with an AI consultant?
It depends on the shape of the engagement, so I keep the shapes explicit. Advisory and fractional-CTO work runs on a monthly retainer sized by cadence. An architecture sprint or production-readiness audit is fixed-scope, with the price agreed before we start, so there is no meter running. Due diligence is a defined deliverable. What I do not do is open-ended hourly billing: it rewards slowness and you cannot budget around it. Call me, describe the problem, and you will leave the first conversation knowing the engagement shape and what it costs.
Let's get your AI to production.
If the PoC-to-production gap is what you're staring at, it's exactly the work I do. Call me, or find me on LinkedIn.