LLMOps changes the unit of engineering from a single model to a distributed service graph. For teams building agentic systems, the practical question is no longer whether an LLM can answer, but how requests move through planning, retrieval, tool calls, retries, memory, and handoffs. That shifts architecture toward explicit boundaries, testable interfaces, and failure isolation rather than a โdrop-in modelโ mindset.
Operational visibility also needs to mature. Standard dashboards are not enough when outputs are nondeterministic and the same prompt can produce different paths. Teams need logging that captures decision points, prompt variants, tool usage, and intermediate state so they can detect anomalies, not just inspect traces after the fact. The trade-off is clear: deeper telemetry improves diagnosability, but it also increases storage, privacy, and governance burden.
Cost discipline becomes a design constraint, not a finance exercise. Token economics, memory pressure, caching strategy, and compute placement all affect reliability and spend. In practice, that means profiling workloads before optimizing kernels, deciding what belongs in RAM versus cache, and aligning model use with the minimum viable level of reasoning. Otherwise, infrastructure choices can quietly erase any productivity gains from generative AI.
The long-term maintainability issue is data and evaluation fidelity. Agent workflows that mirror human processes too closely may encode unnecessary complexity, while static test environments can become obsolete as systems and user behavior change. The most durable implementations will likely combine distributed-systems thinking, continuously refreshed data pipelines, and evaluation loops that reflect real operational drift.
Transcript
This transcript was created with the help of AI and has been lightly edited for clarity. 00.00: All right, so today we have Abi Aryan. She is the author of the OโReilly book on LLMOps as well as the founder of Abide AI. So, Abi, welcome to the podcast.ย 00.19: Thank you so much, Ben. 00.21: All right. Letโs start with the book, which I confess, I just cracked open: LLMOps. People probably listening to this have heard of MLOps. So at a high level, the models have changed: Theyโre bigger, theyโre generative, and so on and so forth. So since youโve written this book, have you seen a wider acceptance of the need for LLMOps?ย 00.51: I think more recently there are more infrastructure companies. So there was a conference happening recently, and there was this sort of perception or messaging across the conference, which was โMLOps is dead.โ Although I donโt agree with that. Thereโs a big difference that companies have started to pick up on more recently, as the infrastructure around the space has sort of started to improve. Theyโre starting to realize how different the pipelines were that people managed and grew, especially for the older companies like Snorkel that were in this space for years and years before large language models came in. The way they were handling data pipelinesโand even the observability platforms that weโre seeing todayโhave changed tremendously. 01.40: What about, Abi, the general.ย .ย .? We donโt have to go into specific tools, but we can if you want. But, you know, if you look at the old MLOps person and then fast-forward, this person is now an LLMOps person. So on a day-to-day basis [has] their suite of tools changed?ย 02.01: Massively. I think for an MLOps person, the focus was very much around โThis is my model. How do I containerize my model, and how do I put it in production?โ That was the entire problem and, you know, most of the work was around โCan I containerize it? What are the best practices around how I arrange my repository? Are we using templates?โ Drawbacks happened, but not as much because most of the time the stuff was tested and there was not too much indeterministic behavior within the models itself. Now that has changed. 02.38: [For] most of the LLMOps engineers, the biggest job right now is doing FinOps really, which is controlling the cost because the models are massive. The second thing, which has been a big difference, is we have shifted from โHow can we build systems?โ to โHow can we build systems that can perform, and not just perform technically but perform behaviorally as well?โ: โWhat is the cost of the model? But also what is the latency? And see whatโs the throughput looking like? How are we managing the memory across different tasks?โ The problem has really shifted when we talk about it.ย .ย . So a lot of focus for MLOps was โLetโs create fantastic dashboards that can do everything.โ Right now itโs no matter which dashboard you create, the monitoring is really very dynamic. 03.32: Yeah, yeah. As you were talking there, you know, I started thinking, yeah, of course, obviously now the inference is essentially a distributed computing problem, right? So that was not the case before. Now you have different phases even of the computation during inference, so you have the prefill phase and the decode phase. And then you might need different setups for those.ย So anecdotally, Abi, did the people who were MLOps people successfully migrate themselves? Were they able to upskill themselves to become LLMOps engineers? 04.14: I know a couple of friends who were MLOps engineers. They were teaching MLOps as wellโDatabricks folks, MVPs. And they were now transitioning to LLMOps. But the way they started is they started focusing very much on, โCan you do evals for these models? They werenโt really dealing with the infrastructure side of it yet. And that was their slow transition. And right now theyโre very much at that point where theyโre thinking, โOK, can we make it easy to just catch these problems within the modelโinferencing itself?โ 04.49: A lot of other problems still stay unsolved. Then the other side, which was like a lot of software engineers who entered the field and became AI engineers, they have a much easier transition because software.ย .ย . The way I look at large language models is not just as another machine learning model but literally like software 3.0 in that way, which is itโs an end-to-end system that will run independently. Now, the model isnโt just something you plug in. The model is the product tree. So for those people, most software is built around these ideas, which is, you know, we need a strong cohesion. We need low coupling. We need to think about โHow are we doing microservices, how the communication happens between different tools that weโre using, how are we calling up our endpoints, how are we securing our endpoints?โ Those questions come easier. So the system design side of things comes easier to people who work in traditional software engineering. So the transition has been a little bit easier for them as compared to people who were traditionally like MLOps engineers. 05.59: And hopefully your book will help some of these MLOps people upskill themselves into this new world. Letโs pivot quickly to agents. Obviously itโs a buzzword. Just like anything in the space, it means different things to different teams. So how do you distinguish agentic systems yourself? 06.24: There are two words in the space. One is agents; one is agent workflows. Basically agents are the components really. Or you can call them the model itself, but theyโre trying to figure out what you meant, even if you forgot to tell them. Thatโs the core work of an agent. And the work of a workflow or the workflow of an agentic system, if you want to call it, is to tell these agents what to actually do. So one is responsible for execution; the other is responsible for the planning side of things. 07.02: I think sometimes when tech journalists write about these things, the general public gets the notion that thereโs this monolithic model that does everything. But the reality is, most teams are moving away from that design as you, as you describe. So they have an agent that acts as an orchestrator or planner and then parcels out the different steps or tasks needed, and then maybe reassembles in the end, right? 07.42: Coming back to your point, itโs now less of a problem of machine learning. Itโs, again, more like a distributed systems problem because we have multiple agents. Some of these agents will have more loadโthey will be the frontend agents, which are communicating to a lot of people. Obviously, on the GPUs, these need more distribution. 08.02: And when it comes to the other agents that may not be used as much, they can be provisioned based on โThis is the need, and this is the availability that we have.โ So all of that provisioning again is a problem. The communication is a problem. Setting up tests across different tasks itself within an entire workflow, now that becomes a problem, which is where a lot of people are trying to implement context engineering. But itโs a very complicated problem to solve. 08.31: And then, Abi, thereโs also the problem of compounding reliability. Letโs say, for example, you have an agentic workflow where one agent passes off to another agent and yet to another third agent. Each agent may have a certain amount of reliability, but it compounds over time. So it compounds across this pipeline, which makes it more challenging.ย 09.02: And thatโs where thereโs a lot of research work going on in the space. Itโs an idea that Iโve talked about in the book as well. At that point when I was writing the book, especially chapter four, in which a lot of these were described, most of the companies right now are [using] monolithic architecture, but itโs not going to be able to sustain as we go towards application. We have to go towards a microservices architecture. And the moment we go towards microservices architecture, there are a lot of problems. One will be the hardware problem. The other is consensus building, which is.ย .ย . Letโs say you have three different agents spread across three different nodes, which would be running very differently. Letโs say one is running on an edge one hundred; one is running on something else. How can we achieve consensus if even one of the nodes ends up winning? So thatโs open research work [where] people are trying to figure out, โCan we achieve consensus in agents based on whatever answer the majority is giving, or how do we really think about it?โ It should be set up at a threshold at which, if itโs beyond this threshold, then you know, this perfectly works. One of the frameworks that is trying to work in this space is called MassGenโtheyโre working on the research side of solving this problem itself in terms of the tool itself. 10.31: By the way, even back in the microservices days in software architecture, obviously people went overboard too. So I think that, as with any of these new things, thereโs a bit of trial and error that you have to go through. And the better you can test your systems and have a setup where you can reproduce and try different things, the better off you are, because many times your first stab at designing your system may not be the right one. Right?ย 11.08: Yeah. And Iโll give you two examples of this. So AI companies tried to use a lot of agentic frameworks. You know people have used Crew; people have used n8n, theyโve used.ย .ย . 11.25: Oh, I hate those! Not I hate.ย .ย . Sorry. Sorry, my friends and crew. 11.30: And 90% of the people working in this space seriously have already made that transition, which is โWe are going to write it ourselves. The same happened for evaluation: There were a lot of evaluation tools out there. What they were doing on the surface is literally just tracing, and tracing wasnโt really solving the problemโit was just a beautiful dashboard that doesnโt really serve much purpose. Maybe for the business teams. But at least for the ML engineers who are supposed to debug these problems and, you know, optimize these systems, essentially, it was not giving much other than โWhat is the error response that weโre getting to everything?โ 12.08: So again, for that one as well, most of the companies have developed their own evaluation frameworks in-house, as of now. The people who are just starting out, obviously theyโve done. But most of the companies that started working with large language models in 2023, theyโve tried every tool out there in 2023, 2024. And right now more and more people are staying away from the frameworks and launching and everything. People have understood that most of the frameworks in this space are not superreliable. 12.41: And [are] also, honestly, a bit bloated. They come with too many things that you donโt need in many ways.ย .ย . 12:54: Security loopholes as well. So for example, like I reported one of the security loopholes with LangChain as well, with LangSmith back in 2024. So those things obviously get reported by people [and] get worked on, but the companies arenโt really proactively working on closing those security loopholes. 13.15: Two open source projects that I like that are not specifically agentic are DSPy and BAML. Wanted to give them a shout out. So this point Iโm about to make, thereโs no easy, clear-cut answer. But one thing I noticed, Abi, is that people will do the following, right? Iโm going to take something we do, and Iโm going to build agents to do the same thing. But the way we do things is I have aโIโm just making this upโI have a project manager and then I have a designer, I have role B, role C, and then thereโs certain emails being exchanged. So then the first step is โLetโs replicate not just the roles but kind of the exchange and communication.โ And sometimes that actually increases the complexity of the design of your system because maybe you donโt need to do it the way the humans do it. Right? Maybe if you go to automation and agents, you donโt have to over-anthropomorphize your workflow. Right. So what do you think about this observation? 14.31: A very interesting analogy Iโll give you is people are trying to replicate intelligence without understanding what intelligence is. The same for consciousness. Everybody wants to replicate and create consciousness without understanding consciousness. So the same is happening with this as well, which is we are trying to replicate a human workflow without really understanding how humans work. 14.55: And sometimes humans may not be the most efficient thing. Like they exchange five emails to arrive at something.ย 15.04: And humans are never context defined. And in a very limiting sense. Even if somebodyโs job is to do editing, theyโre not just doing editing. They are looking at the flow. They are looking for a lot of things which you canโt really define. Obviously you can over a period of time, but it needs a lot of observation to understand. And that skill also depends on who the person is. Different people have different skills as well. Most of the agentic systems right now, theyโre just glorified Zapier IFTTT routines. Thatโs the way I look at them right now. The if recipes: If this, then that. 15.48: Yeah, yeah. Robotic process automation I guess is what people call it. The other thing that people I donโt think understand just reading the popular tech press is that agents have levels of autonomy, right? Most teams donโt actually build an agent and unleash it full autonomous from day one. I mean, I guess the analogy would be in self-driving cars: They have different levels of automation. Most enterprise AI teams realize that with agents, you have to kind of treat them that way too, depending on the complexity and the importance of the workflow.ย So you go first very much a human is involved and then less and less human over time as you develop confidence in the agent. But I think itโs not good practice to just kind of let an agent run wild. Especially right now.ย 16.56: Itโs not, because whoโs the person answering if the agent goes wrong? And thatโs a question that has come up often. So this is the work that weโre doing at Abide really, which is trying to create a decision layer on top of the knowledge retrieval layer. 17.07: Most of the agents which are built using just large language models.ย .ย . LLMsโI think people need to understand this partโare fantastic at knowledge retrieval, but they do not know how to make decisions. If you think agents are independent decision makers and they can figure things out, no, they cannot figure things out. They can look at the database and try to do something. Now, what they do may or may not be what you like, no matter how many rules you define across that. So what we really need to develop is some sort of symbolic language around how these agents are working, which is more like trying to give them a model of the world around โWhat is the cause and effect, with all of these decisions that youโre making? How do we prioritize one decision where the.ย .ย .? What was the reasoning behind that so that entire decision making reasoning here has been the missing part?โ 18.02: You brought up the topic of observability. Thereโs two schools of thought here as far as agentic observability. The first one is we donโt need new tools. We have the tools. We just have to apply [them] to agents. And then the second, of course, is this is a new situation. So now we need to be able to do more.ย .ย . The observability tools have to be more capable because weโre dealing with nondeterministic systems. And so maybe we need to capture more information along the way. Chains of decision, reasoning, traceability, and so on and so forth. Where do you fall in this kind of spectrum of we donโt need new tools or we need new tools?ย 18.48: We donโt need new tools, but we certainly need new frameworks, and especially a new way of thinking. Observability in the MLOps worldโfantastic; it was just about tools. Now, people have to stop thinking about observability as just visibility into the system and start thinking of it as an anomaly detection problem. And that was something Iโd written in the book as well. Now itโs no longer about โCan I see what my token length is?โ No, thatโs not enough. You have to look for anomalies at every single part of the layer across a lot of metrics. 19.24: So your position is we can use the existing tools. We may have to log more things.ย 19.33: We may have to log more things, and then start building simple ML models to be able to do anomaly detection. Think of managing any machine, any LLM model, any agent as really like a fraud detection pipeline. So every single time youโre looking for โWhat are the simplest signs of fraud?โ And that can happen across various factors. But we need more logging. And again you donโt need external tools for that. You can set up your own loggers as well. Most of the people I know have been setting up their own loggers within their companies. So you can simply use telemetry to be able to a.) define a set and use the general logs, and b.) be able to define your own custom logs as well, depending on your agent pipeline itself. You can define โThis is what itโs trying to doโ and log more things across those things, and then start building small machine learning models to look for whatโs going on over there. 20.36: So what is the state of โWhere we are? How many teams are doing this?โ 20.42: Very few. Very, very few. Maybe just the top bits. The ones who are doing reinforcement learning training and using RL environments, because thatโs where theyโre getting their data to do RL. But people who are not using RL to be able to retrain their model, theyโre not really doing much of this part; theyโre still depending very much on external accounts. 21.12: Iโll get back to RL in a second. But one topic you raised when you pointed out the transition from MLOps to LLMOps was the importance of FinOps, which is, for our listeners, basically managing your cloud computing costsโor in this case, increasingly mastering token economics. Because basically, itโs one of these things that I think can bite you. For example, the first time you use Claude Code, you go, โOh, man, this tool is powerful.โ And then boom, you get an email with a bill. I see, thatโs why itโs powerful. And you multiply that across the board to teams who are starting to maybe deploy some of these things. And you see the importance of FinOps. So where are we, Abi, as far as tooling for FinOps in the age of generative AI and also the practice of FinOps in the age of generative AI?ย 22.19: Less than 5%, maybe even 2% of the way there. 22:24: Really? But obviously everyoneโs aware of it, right? Because at some point, when you deploy, you become aware.ย 22.33: Not enough people. A lot of people just think about FinOps as cloud, basically the cloud cost. And there are different kinds of costs in the cloud. One of the things people are not doing enough is not profiling their models properly, which is [determining] โWhere are the costs really coming from? Our modelsโ compute power? Are they taking too much RAM? 22.58: Or are we using reasoning when we donโt need it? 23.00: Exactly. Now thatโs a problem we solve very differently. Thatโs where yes, you can do kernel fusion. Define your own custom kernels. Right now thereโs a massive number of people who think we need to rewrite kernels for everything. Itโs only going to solve one problem, which is the compute-bound problem. But itโs not going to solve the memory-bound problem. Your data engineering pipelines arenโt whatโs going to solve your memory-bound problems. And thatโs where most of the focus is missing. Iโve mentioned it in the book as well: Data engineering is the foundation of first being able to solve the problems. And then we moved to the compute-bound problems. Do not start optimizing the kernels over there. And then the third part would be the communication-bound problem, which is โHow do we make these GPUs talk smarter with each other? How do we figure out the agent consensus and all of those problems?โ Now thatโs a communication problem. And thatโs what happens when there are different levels of bandwidth. Everybodyโs dealing with the internet bandwidth as well, the kind of serving speed as well, different kinds of cost and every kind of transitioning from one node to another. If weโre not really hosting our own infrastructure, then thatโs a different problem, because it depends on โWhich server do you get assigned your GPUs on again?โ 24.20: Yeah, yeah, yeah. I want to give a shout out to RayโIโm an advisor to Anyscaleโbecause Ray basically is built for these sorts of pipelines because it can do fine-grained utilization and help you decide between CPU and GPU. And just generally, you donโt think that the teams are taking token economics seriously? I guess not. How many people have I heard talking about caching, for example? Because if itโs a prompt that [has been] answered before, why do you have to go through it again?ย 25.07: I think plenty of people have started implementing KV caching, but they donโt really know.ย .ย . Again, one of the questions people donโt understand is โHow much do we need to store in the memory itself, and how much do we need to store in the cache?โ which is the big memory question. So thatโs the one I donโt think people are able to solve. A lot of people are storing too much stuff in the cache that should actually be stored in the RAM itself, in the memory. And there are generalist applications that donโt really understand that this agent doesnโt really need access to the memory. Thereโs no point. Itโs just lost in the throughput really. So I think the problem isnโt really caching. The problem is that differentiation of understanding for people. 25.55: Yeah, yeah, I just threw that out as one element. Because obviously thereโs many, many things to mastering token economics. So you, you brought up reinforcement learning. A few years ago, obviously people got really into โLetโs do fine-tuning.โ But then they quickly realized.ย .ย . And actually fine-tuning became easy because basically there became so many services where you can just focus on labeled data. You upload your labeled data, boom, come back from lunch, you have a fine-tuned model. But then people realize that โI fine-tuned, but the model that results isnโt really as good as my fine-tuning data.โ And then obviously RAG and context engineering came into the picture. Now it seems like more people are again talking about reinforcement learning, but in the context of LLMs. And thereโs a lot of libraries, many of them built on Ray, for example. But it seems like whatโs missing, Abi, is that fine-tuning got to the point where I can sit down a domain expert and say, โProduce labeled data.โ And basically the domain expert is a first-class participant in fine-tuning. As best I can tell, for reinforcement learning, the tools arenโt there yet. The UX hasnโt been figured out in order to bring in the domain experts as the first-class citizen in the reinforcement learning processโwhich they need to be because a lot of the stuff really resides in their brain.ย 27.45: The big problem here, and very, very much to the point of what you pointed out, is the tools arenโt really there. And one very specific thing I can tell you is most of the reinforcement learning environments that youโre seeing are static environments. Agents are not learning statically. They are learning dynamically. If your RL environment cannot adapt dynamically, which basically in 2018, 2019, emerged as the OpenAI Gym and a lot of reinforcement learning libraries were coming out. 28.18: There is a line of work called curriculum learning, which is basically adapting your modelโs difficulty to the results itself. So basically now that can be used in reinforcement learning, but Iโve not seen any practical implementation of using curriculum learning for reinforcement learning environments. So people create these environmentsโfantastic. They work well for a little bit of time, and then they become useless. So thatโs where even OpenAI, Anthropic, those companies are struggling as well. Theyโve paid heavily in contracts, which are yearlong contracts to say, โCan you build this vertical environment? Can you build that vertical environment?โ and that works fantastically But once the model learns on it, then thereโs nothing else to learn. And then you go back into the question of, โIs this data fresh? Is this adaptive with the world?โ And it becomes the same RAG problem over again. 29.18: So maybe the problem is with RL itself. Maybe maybe we need a different paradigm. Itโs just too hard.ย Let me close by looking to the future. The first thing isโthe space is moving so hard, this might be an impossible question to ask, but if you look at, letโs say, 6 to 18 months, what are some things in the research domain that you think are not being talked enough about that might produce enough practical utility that we will start hearing about them in 6 to 12, 6 to 18 months? 29.55: One is how to profile your machine learning models, like the entire systems end-to-end. A lot of people do not understand them as systems, but only as models. So thatโs one thing which will make a massive amount of difference. There are a lot of AI engineers today, but we donโt have enough system design engineers. 30.16: This is something that Ion Stoica at Sky Computing Lab has been giving keynotes about. Yeah. Interesting.ย 30.23: The second part is.ย .ย . Iโm optimistic about seeing curriculum learning applied to reinforcement learning as well, where our RL environments can adapt in real time so when we train agents on them, they are dynamically adapting as well. Thatโs also [some] of the work being done by labs like Circana, which are working in artificial labs, artificial light frame, all of that stuffโevolution of any kind of machine learning model accuracy. 30.57: The third thing where I feel like the communities are falling behind massively is on the data engineering side. Thatโs where we have massive gains to get. 31.09: So on the data engineering side, Iโm happy to say that I advise several companies in the space that are completely focused on tools for these new workloads and these new data types.ย Last question for our listeners: What mindset shift or what skill do they need to pick up in order to position themselves in their career for the next 18 to 24 months? 31.40: For anybody whoโs an AI engineer, a machine learning engineer, an LLMOps engineer, or an MLOps engineer, first learn how to profile your models. Start picking up Ray very quickly as a tool to just get started on, to see how distributed systems work. You can pick the LLM if you want, but start understanding distributed systems first. And once you start understanding those systems, then start looking back into the models itself.ย32.11: And with that, thank you, Abi.
Enjoyed this article? Sign up for our newsletter to receive regular insights and stay connected.

