All episodes
S1 · Ep 6Ontology Engineering

An ontology fits on a Post-it note

Casey Hart is a rare thing: an actual ontologist. A philosophy PhD who answered a job ad from Cycorp and spent a decade building knowledge for machines under Doug Lenat and then at Olive, Amazon, Gro Intelligence, and Ford, he spends this episode deflating the word everyone is suddenly selling. An ontology, he argues, is just a summary of what your business cares about and how those things relate — you can start one on a Post-it note. He separates the machine-learning "system one" from the deterministic "system two" that ontologies supply, makes the case for a hybrid, and walks through building one from the ground up: taxonomies, relationships, turtle files and triple stores — or just the metadata, so you get value before migrating a single row. Along the way: why "hallucination" flatters a text generator doing exactly what it was built to do, the open-world versus closed-world assumption, and why vibe-coding an ontology out of an LLM is a fine way in but not a finished asset.

Casey HartBio ↓
Ontologist & Founder · Casey Hart Consulting
Watch

Watch the conversation

The speaker

Casey Hart

Ontologist & Founder · Casey Hart Consulting

A philosopher-turned-ontologist and founder of Casey Hart Consulting, where he helps organizations understand what ontologies are, what they can do, and how to actually build them. He fell into the field almost by accident — a philosophy PhD from the University of Wisconsin–Madison who answered a posting from Cycorp and spent three years building common-sense knowledge in CycL under Doug Lenat, before developing core ontologies at Olive, Amazon, and Gro Intelligence, and today at Ford. His academic background is in formal and social epistemology, philosophy of science, and metaphysics; he runs the YouTube channel "Ontology Explained" and co-hosts the podcast "Philosophy, Programs, and Prompts," and works most in healthcare, automotive, and energy.

Episode evaluation

What to do with this episode

Four clear reads — who should act, and how urgently.

01For buildersTEST

Ask an LLM to draft your first ontology

Feed a dataset to an LLM and have it generate a starter ontology — it's a relatively painless way to feel out how the model maps to your data. Just don't ship the rat's nest: valid turtle that compiles is no more "good" than Python that compiles, so keep a qualified hand between the draft and production.

02For data & AI leadsUSE

Start with metadata, not migration

You don't have to move every row into a graph. Model the metadata first — what data you have, where it lives, what the columns actually mean — and you get analyst-time savings and duplicate-column cleanup before any heavy ETL.

03For enterprise buyersWATCH

Don't build one for a job that never changes

If your data always looks the same and you do the same thing with it every time, a plain program beats an ontology. The payoff is reuse and future options — commission one only when you'll actually build on it, and treat it as an ongoing effort, not a deliverable.

04Bottom lineSHIP

Say true things, and start on a Post-it note

An ontology is a summary of what your business cares about and how those things relate. Write down as many true things as you can, leave room to get more specific later, and grow it as it earns its keep. Responsible use of AI — not "responsible AI."

Show notes

We discuss

  • 01How a philosophy PhD becomes an ontologist — and why "thinking fast / thinking slow" makes philosophers a natural fit for grounding machine reasoning.
  • 02Why "hallucination" flatters a text generator that's doing exactly what it was built to do — and why "generation" is the more honest word.
  • 03Cyc's ultra-expressive representation vs the "three-word sentences" of OWL/RDF — n-ary relations, reification, and whether RDF-star (1.2) actually adds expressivity.
  • 04The airplane-manufacturer use case: an ontology as a map of what data you have and where it lives, before you migrate a single row.
  • 05Building one from the ground up — taxonomies, relationships, turtle files, triple stores — and why you still need someone qualified, not just an LLM's rat's nest.
  • 06Open-world vs closed-world: why "I didn't record Casey's SSN" doesn't mean he doesn't have one — and when you do want to close the world.
  • 07Start small: an ontology on a Post-it note, "data therapy," and saying as many true things as you can.
  • 08Who owns the ontology in an org — data engineering, product, or wherever the first use case lands — and the responsible way to use an LLM to draft one.
Reference

Transcript

VivekAll right, welcome back to ContextOps. Hello, Casey. How are you?

CaseyI'm excellent. Pleased to be here.

VivekThank you so much for taking the time for doing this, Casey. Really, really appreciate this. Right. Now, for the audience context, Casey Hart is an actual ontologist. Can you guys believe that? For the longest time, this title was something that I think Palantir had a copyright on or something like that. But ontologists actually exist. Casey is a real one. and today we will dive deeper into figuring out what actually is an ontology, apart from everything that you have seen on LinkedIn and Twitter and Reddit and YouTube and so on. Right. But before we get into this, Casey, right, how does somebody with a PhD in philosophy ends up working on what some of us call as foundation for generative AI. Like how does that happen?

CaseyYeah. I mean, the way I got in there had nothing to do with generative AI. It was a little bit pre-generative AI. But the path to becoming an ontologist is wide and varied. It's something I talk about on my YouTube channel and with other academic philosophers who are maybe thinking about their way out. The way my particular journey happened was I was getting my PhD in philosophy, which means you want to teach it and be a professor almost certainly unless you're like getting into law or something, but then you probably just get your MA and then go there. And philosophers go to philjobs.com to find their job.

And so I went there and you apply to every area of, you know, AOS or AOC that you can squint and pretend like you qualify for. And one of those qualifications was for this random company called Cycorp that put a posting up, and I was like, that's not academic, but whatever, I'm applying for all of that stuff. And I got an interview at Cycorp with Doug Lenat and got offered a position. It was between that and some post-docs. And I was like, well, this is a job that I'll view as a post-doc that pays me a little more money. And that we can sort of settle down.

My wife and I had a two-year-old at the time and a... negative one month old at the time. like, we don't want to be traveling a bunch and like immediately reapplying for new academic postings. And so I found myself as an ontologist thought I would be there for, you know, a year or two. And then like my, my, my deep and detailed plan was I'll reach out to some liberal arts applications, right? And I can tell those business deans that not only can I teach philosophy, but I know how to get a real paycheck using it. And that'll make me irresistible. And before that happened, was like, Wait, this ontology thing is super fun.

I get to, like, it scratches a lot of the philosophical itches. I get to do some, you know, like learn new things, formalize stuff, think carefully about logic and inferences. And so the, I just found myself getting, getting farther into the career. Yeah.

VivekThat's such a believable story.

CaseyHehehehe, both. Good, because I made it up. So I'm glad you believed it. Yeah. I was just trying out new things every time.

VivekYeah, No, no, no, it all adds up, right? You know, a philosophy PhD looking for a way out of a postdoc. Yeah, I mean it makes sense, right?

CaseyBut I will say like, it's not, know, from that, it's like, it's a weird journey or whatever. And they're like, I've met jazz musicians who got into ontology, all these sorts of things. I would say philosophy actually makes a lot of sense as a way in, right? So when you are in ontology, a bunch of ways to conceive of it, but one of them is like, it's the, you know, if you take that like left brain, right brain, Kahneman and Tversky view of psychology, where we have like this thinking fast. We're just sort of like gut reaction. That's kind of like machine learning ish. And then you've got the thinking slow, like logical and more methodical approach.

That's where good old fashioned AI, that's where ontology sits is building like a model that can ground those sorts of reasonings. And that's the kind of stuff that I think philosophers are pretty good at. Like let's stop, let's reflect on what our assumptions are. Let's formalize those assumptions. Let's see how they fit into arguments that would generate the sorts of inferences and conclusions that we want. Right? That kind of, you know, combination of like carefulness and a little bit of like creativity and mathematical formulations all sort of suits itself pretty nicely, I think, to this field.

VivekI agree, because Josh here himself has a strong philosophy background. one of the ways we describe

JoshuaNo, I mean that's not true, not philosophy, but maths. And it has overlaps, but it's not philosophy, yeah.

VivekI stand corrected. I thought it was potato potato, but okay. Maths. Sure.

JoshuaWell, it's like, you know, if you were a doctorate engineering student, they would be like, maths is about calculations. And that's not what we studied. I mean, yeah, we had some calculations. But we had things like, you know, why is zero plus zero zero? And that's a stupid question. these are the stupid questions that we had to prove. And that was what maths was about. So in that sense, yeah, it has some overlaps. I think it used to be taught together and now it's different for good reason. But yeah, I mean, I'm not a philosopher, but I like math. I don't regret. that journey. That was my fun times.

VivekYeah, So Casey, one question which you obviously have been building this, building ontologies before Transformers became popular and before GPT three point five and so on, right? But when these models sort of came around and they become they became like extremely important and critical to a lot of people, right? Wha what was that point when you realize that, nah, this ain't it? Like when did Ha ha. you realize that? Like, no. This still needs a lot of exactly what you said, right? It needs a system two. This is only a system one. Right. Yeah.

CaseyYeah, So some of it, was never that sold on it being everything, right? Like when you say this ain't it, which is a sentiment I agree with, I don't want to overstate it and be like, oh, I think LLMs, generative AI are worthless tools. They don't do anything for us. They're super cool. They do lots of awesome things. A lot of the discoveries and sort of the types of answers that you can get from LLMs, are like surprising and understanding why we get those kinds of results is important and good both in computer science and psychology and all sorts of things. So there you go.

Now I've said the positive things. Now I can go back to negative things. Why would I be negative about it? One is a contrarian spirit that I have perhaps. But just for the longest time, know, LLMs still fit broadly under the machine learning, you know, umbrella of things, right? And so from, you know, as someone who learned a lot about artificial intelligence And so, as someone who learned a lot about artificial intelligence under Doug Lenat, right brain sort of stuff. So in that sense, there was nothing new here. Right. But the other piece for me is just when you're thinking about like fundamentally what is going on, we're we've got a bunch of training data.

We are building up a model that synthesizes, compresses the information in that training data, right? And this training data is a bunch of, you know, it depends on the model that you're looking at, but like text from the internet. And there's a bunch of like, that text correlates with a bunch of important knowledge, that's true, right? But as we're compressing that down, it's going, it would be, it's just not possible to think that's got like a total understanding of what the world is like, right? And it's a little bit of a pejorative. to say that all LLMs are doing is just guessing the next word.

But there's something importantly true about like that is the kind of thing that's going on. It turns out that's a really powerful technique and that you can look at bigger contexts and those sorts of things. But fundamentally, it's just trying to be a pattern predictor about the kind of thing that would make a sensible, like the kind of sentence you would see in the training data. And of course, that's not inference. That's not... knowing facts about the world and then reciting those facts to you. So you're gonna end up with these hallucinations, other sorts of things. I hate calling them hallucinations, but I think the term is stuck,

VivekYeah. We we also I think we internally use the word generation more than hallucination. Right? That it's generating stuff. Just that it aligns with your point of view at times and it's at that time it's not a hallucination. And when it does not align with your point of view, it's a hallucination or, you know, something of that sort, right? It triggers that reaction from from people.

CaseyRight. The thing that bothers me about the hallucination is like, I almost just want to call it a bug, right? Like we have a text generator and it's generating bad text. That's a bug, right? But when we make it a hallucination, not only do we sort of give it an excuse or something, right? We have like built it up to be more like psychologically sophisticated than it is, right? Like we don't, we don't do that for other sorts of like, You know, if you cut the brake lines to my car or something and like I'd run into the garage instead of coming to a stop and the manufacturer was something, you know, said something like, the car made an educated decision that stopping was unwise at that point.

I'd be like, no, like, car wasn't thinking right. Like you just gave me a bad product. And yet when LLMs screw things up, we like make them seem like they're even smarter. this is so sophisticated that it's messing up. Like that's not. That's not right.

VivekYeah, no, I I I totally agree. The the other version of this, which also gets discussed is when you call it a bug, it also means you can fix it or it's fixable, right? But the reality is you try the same questi same request and you might actually get a better output, right? Or the other way of looking at it is which is if it's it makes it makes it sound like it has some disease and that you can fix it. Or there's a healthier version of this LLM instance running somewhere out there. So let's go chase that healthy version, right? Let's get the LLM with less sugar and l or less carbs or whatever that is, right? and that really did that and that's that's not reality, right? every model, every LLM will have the same core traits, which is generate when you can. Right. yeah.

CaseyGood, it's doing, I see that, I appreciate that. So instead of calling it a bug, which suggests that we remedy it, and maybe there are ways to change the model or the software that sits on top of the model to make more generally appealing outputs. But fundamentally, the thing wasn't doing anything wrong. It was performing exactly as intended, which is making text in accordance with the vector spaces and the various rules and weights and. random number generators that you assigned to them. Yeah, so it was doing what they asked it to do. It just turns out that we don't like all the things that we asked it to do. Yeah.

VivekThat's it. That's that's that's pretty much it. Alright. So see going back to your your first stint as an ontologist, right? Where you were that work was pretty much in their own proprietary representation language, which was which is which I think you've mentioned is to be very different from the OWL and the RDF world where most of the conversation today is happening in. Right. how was that experience in terms of learning which is not today happening let's say in the RDF world, or what is it that you can carry from there over here?

CaseyYeah. great. So coming from philosophy, right? And, you know, I did some formal stuff, you getting back to talking about Josh and math and philosophy and how these things tied together. Like I did Bayesian epistemology, you know, I did some, uh, you know, various logics. Uh, so I was used to some, you know, symbolism and formalization and computation, but I was not a programmer by any stretch. And when I got to Cycorp, they had their language, which was CycL. So it is C-Y-C, as in encyclopedia. That is Cycorp. The company sounds so evil, right? Like it's, it would fit in the, you know, sky net or whatever that you're doing in the Terminator, but CycL is like a Lisp variant out there.

So I didn't have any. basis to say like, is this more or less complicated than it ought to be? I was just learning the representation and it didn't seem way harder than any of the like quantified modal logics that I had been dealing with. So it was fine. And had some really nice programmers there who took me under their wing and, you know, thanks to Emma and Andrew and Andrew for putting up with my stupid questions about how to write Lisp functions. And so I learned that stuff. The big differentiator between, And you can think of OWL/RDF as having some short, lispy kinds of constructions too.

So those are not all that different. The biggest difference was at Cyc: they wanted like a very robust expressive language. And so you could have predicates or properties that related as many objects as you liked. Right. So in OWL/RDF, you have basically two-place relations. You're making a graph, right? So our knowledge graph has nodes connected by edges, right? So I can say like Casey ate pizza. So Casey and pizza are nodes in our graph and eight is a predicate or a property or an edge that connects those two types of nodes. but of course we can have really complicated, properties, right? Like we could have, you know, Vivek and Casey and Josh.

do a podcast together. And so I could make the do a podcast together where I could make participate in event type, know, event type podcast participant, and I can have a list of as many objects as I want. So you can make those relationships crazy, right? And now you're not thinking about a two dimensional or maybe even three dimensional graph. You're in like N dimensional space and it would look really crazy if you were trying to represent it there. But basically, I really liked that as a basis for saying, take any simple like, as complicated a sentence as you want, and you should be able to represent that in your ontology, right?

All the facts about the world, we should be able to put that into our ontology. And Cycorp in particular had the mandate of we want to get common sense reasoning into the knowledge base. And that means we want to represent facts in as general a way as possible. And so instead of representing some facts about I don't know, like my tennis racket fits inside my tennis bag, right? I want to say, well, there's a general principle here that is smaller objects can fit inside larger objects or something like that, or larger objects can't fit inside objects that are smaller than them. And this is going to be used for like whether my tennis racket can fit into my backpack as opposed to my tennis bag or whether the horse can fit into my pathfinder rather than, you know, fit into the horse trailer or whatever you like.

so you, you have this like stating general rules about reality and thinking being not constrained and trying to fit that into smaller constructions was actually really helpful. When I went to OWL/RDF, that was, you know, partly at the time was like, wait, I think I like ontology. I'm not going to just leave and go back to academia or something like that. so let's see how everybody else are doing, are doing these things. and At Cyc, we always had this insult that we levied at people who were doing OWL and RDF saying like, they only get to say things in like three word sentences, but how ridiculous is that?

Like the world's more complicated than something you could just say in three word sentences. And I got good at giving that line in sales pitches, but I was like, I don't know if that's a good approximation of what's going going on in OWL/RDF or not. And so then I went and played around with it myself and learned how to do stuff. And in some way it's way simpler. So it's easier for me to express stuff. There are times when I'm upset that like, I wanted to say this like more fine-grained things. And now I have to find like tricks to avoid saying the more complicated thing and then, you know, compressing it down.

But it's not too bad. And there are constructions, like you have list objects in RDF. So if you really want, you can sort of secretly make predicates as long as you'd like them, because you can just make subject predicate, and then in the object, you can just put a list and that list can be as long as you like. That's terrible practice. Don't do it most of the time, but sometimes you try and shove a more complex data structure in there for special cases. I am

Joshuaguessing RDF-star is also making things easier.

CaseyI don't think so. So RDF-star, right? Or spec, what is it? 3.2? I forget the 1.2. 1.1. Yeah. So for those listening who may or may not know what that is, say that you want to represent something like the classic example is a marriage, right? So Josh and I get married and then we get divorced. It was a tough first stint, but then we decide to give it a go again and we get remarried. Well, in RDF and OWL, if you say Casey and then later I say Casey married to Josh, those are the same triple. And so they just sort of get compressed. And this is one of the things that people say in favor of like labeled property graphs, which allow you to put special additional data on top of the edges, right?

But the way that you can solve that before in OWL/RDF is instead of saying Casey married to Josh, I can say there's marriage 001 that has participants Casey and Josh, right? And then I can put dates on that. And then I have marriage 002 that has Casey and Josh as participants and that has different date constructions. So instead of just having two edges, I want to distinguish the edges, I reify the events is what we say. So what RDF 1.2, RDF-star does is allow you to express the sentences as triples themselves. So I can say, Casey married to Josh is my sentence. And I can say when that sentence was true.

And then I can say that sentence again, this RDF statement, and then say that was true at different times. But if you could reify the events, right, marriage one and marriage two, or you could take the propositions, the RDF statements, they turn out to be exactly isomorphic structures. So RDF-star, RDF 1.2, I don't think actually it doesn't give you any more expressivity. I was talking to Ora Lassila about this at the recent Knowledge Graph Conference, and he's one of the, he's semantic web Mount Rushmore guy, just awesome dude to talk to. And he is on that W3C committee for making the RDF 1.2 spec. And I feel confident in my assertion, because he feels about the same way, where he's like, yeah, it's isomorphic, you're getting the same structures.

One nice thing about RDF 1.2 is it gives you, it's another place where you can store information and metadata. And so if you wanted to just use those like RDF 1.2 statements to say like who made the assertion or how confident we are in the assertion or something like that, these sort of like, they're not strictly speaking models about the world, but they're meta statements about the data model. That's a good place to hang those things. So maybe it allows me to separate out some of those interests.

JoshuaYeah, I was talking more about the expressivity of the language itself, like to be able to express more and less, which is I think what you were talking about in Cyc.

CaseyYeah, but then because you can reify the events and you could reify the statements if you wanted to in other ways already. Yeah, you could do it on them. So I just don't think it helped you. Yeah It didn't give me any new colors or paint brushes to paint with.

JoshuaSo what would you like to see in RDF or?

CaseyYeah, I I would like some more standard, mostly I think it is on the SPARQL side of things. Like I'd like to have a little bit more control over how queries are run or sequenced in a standard way. In some ways, I'm not missing all that much because you can just build your own wrappers around things and like that's what I've done with my various clients and my personal ontology stuff. But I would love, and I There are some companies who have tried this before, haven't always been super effective with it, but to say like, bind this variable first, right? Here's a long query, go get those ones first because there's only gonna be five of them and that's gonna be easy.

And instead, when I go from triple store to triple store, I have to learn like, is it random how they're picking out these bindings or do they just try and bind the highest up variables or the deepest nested variables first? And like if that stuff was super clear, that would make the SPARQL querying easier and better. And I think improving the query stuff is more important than anything else, because that's what's gonna be powering your APIs and applications and stuff. So if I can make that fast and accurate and standard, that would be awesome.

JoshuaSo more like graph, you know, algorithms walking the graph more easily, being more expressive, giving you more options basically in the query language itself.

CaseyYeah. I mean, for some of the other stuff, right? Like as much as I like the additional expressivity that we got right? As much as I like the additional expressivity that we got at Cyc, that comes with slowness, right? The reason that OWL/RDF

JoshuaSo what happened to Cyc? Sorry. yeah. I think it was like one of those big things, right? Like up there with, was it DBpedia or something like that? Or schema.org, all of these movements that were happening. Is it still around? Is it merged into something? What happened there?

CaseyIt is I think in a coma would be the right way to I think, in a coma, would be the right way to metaphorically describe where I don't think that From the outside without any inside information I've been away from there for a while. It felt like things ran out of steam you know, general economic concerns and other stuff and with Doug's passing that probably didn't didn't help things for, you know, preserving some clients and acquiring some new ones. and it looked like they were hoping to be acquired or something and that stuff wouldn't be as sustainable. That's my, that's my sense of things. yeah, there's, you know, for they had, they had some issues over time, but also they were tremendously successful in other ways to just. Yeah,

Joshuayeah, it was quite popular,

Caseybe doing that since the mid 80s. Yeah. think, like, like, recently, like, decade back. Yeah. And you will find that Cycorp is an incredibly polarizing place. think you'll find lots of people that are like, if we invested more, did a little bit better, like the singularity would be here by now, or like Cyc is the key to it. And then you'll have other people who will tell you, that it was, you know, borderline criminal and abusive enterprise and waste a bunch of days. The truth is always somewhere in between, right? They did some awesome and cool stuff and they gave me a start. I'm forever grateful for that.

And there are lots of amazing things that I wish they could have done a little bit better job at getting out there for long-term business aims and stuff. But it was a fun ride.

JoshuaI guess that was AGI back then,

CaseyYeah. Well, I mean, Right. some of the stuff, like when you want to talk about, you know, there were competing camps, right? Since the, you know, since the origins of AI discussions, you know, like in the fifties and sixties, where you're building computers to model and understand the way the human minds work. And some folks were on the, like the good old fashioned AI expert system side of things and say, I don't care what you do. logical reasoning and computations of what's going on in the brain. And when I say that penguins are birds, but they can't fly, but I also say that all penguins can fly or all birds can fly.

I am doing exception based reasoning. I have these assertions, and Cyc was a model that was trying to get all of those core facts and do exception reasoning and stuff like that. And then in the other camps, right, you have people that are like, nah, it's machine learning is going to get there eventually. And we just need to see when we have enough computing power. And there are the way there was like this race between the expert system size of like, we just need to build up a big enough repository of digital facts that we can leverage to get the kind of reasoning that humans are doing.

other folks are like, we just need to get the processing power and the training data sets to get these kinds of things. And yeah, and some, mean, the debate is still out, right? Some people are super, you know, we talked about whether LLMs are it or not. And, uh, I say no in some respects, right? But there's some people out there like, you guys were dumb because you did not think LLMs could do as much as they could do now. Like if we just get some more processing power and some more training data, then it'll tell us that, know, penguins can't fly, but all birds can fly for exactly the right reasons too. So yeah, but Cycorp was—

JoshuaSo where do you sit? Are you more symbolic, more neural? Where are you?

CaseyI think that the right response is going to be some sort of hybrid response. There's, it depends on what question we're asking. One question is how do these things work, tapping my head while we're on a podcast, right? Like how does the human mind and human brain work? And is it the sort of thing that has parts that are irreducibly computational versus more like machine learning sorts of things? I don't know. Like I don't know enough about psychology to say that sort of thing. I think that's interesting, but then there's the other question about like, what are going to be the best AI answers? And for me, I think that there's going to be some sort of hybrid solution as the right one.

Like machine learning stuff is like obviously really good and really useful, but it's very error prone and it's really computationally expensive. And it's not very transparent. So when we're leveraging a tool like that, that's going to be really valuable in certain cases where we have the right sorts of data, where we want to generate some new insights by taking advantage of computational power. But then there are lots of times when we have to be able to trust the output and I'm not going to trust the output of something. was like, I ran 6 billion computations and generated these vectors. And I give you this answer.

Just take the chemotherapy. No, no, no, thank you. So the right answer for me is going to have to use a bunch of data, machine learning algorithm sorts of things, but it's going to have to have a ground world model and sets of reasons and inferential steps to explain to me why I should be compelled by the output of the machine.

VivekRight. Awesome. So again, I I know you and Josh went at it. but for those who are not, you know, philosophy who are not math grads or philosophy majors or doctorates, right? I think it's important to most people think of ontology as an English word. So I want to go back to absolute basics now for a minute. and for them it's very hard to sort of you know operationalize this word in their GenAI chatbot. So when somebody says your chatbot has an as an ontology problem, it becomes a very difficult statement for them to make sense out of it. Okay. And I'm sure you meet folks like these, I hope you do meet folks like these who ask you such naive questions.

How do you bring them up to speed? Especially when there is so much you know, so much amazing experience that they've had with LLMs. Okay. Where like the LLM just gets me, Okay. man. Come on. Right? Like what's this ontology thing that you're talking about?

CaseyYeah. So, I mean, the pitch depends a lot on who the audience I'm talking to is, like why they're asking me these things, right? If I'm selling consulting services versus whatever. But, and also depends a bit, you when you say they're coming from LLM stuff and then they're asking me these naive questions about ontology, it depends on whether they're saying, like this generative AI is doing such cool stuff for us. Okay. Why do we also need ontology? That's one side. And other people come to me and they're like, I'm sick of LLMs ruining stuff for me. What can you offer? Then there's a different pitch about what, yeah.

what ontology is do. But the, the first pass, I mean, I might do the left brain, right brain sort of thing here, right. Which is that, yes, we want some gut reactions based on a lot of experiences and a lot of your data. What seems like the likely next thing. Right. And so if we're say talking about a medical case or something, I looked at. 15,000 patients who have similar biomarkers as you and 10,000 of them ended up having a hip replacement in their mid-60s, right? That's useful information. That tells me like, I should pay attention to the health of my hip and maybe what supplements should I take or something like that, right?

That's a machine learning-y, LLM-y kind of thing. But I should also be coupling that with things that we understand about the way that the world works. know what human anatomy looks like, right? We have textbooks, we have understanding about how the joints work, what sorts of things promote joint health, what sorts of genetics encode for likelihoods of hip replacements. Now here's where I just don't know that much about the medical of hip replacements, Hips and replacements. Yeah. which is why, of course, you should not pick that example on a podcast, but I ran with it, so here we are, right? So what I think the ontology does is it lays out Here is what we understand the world to be like.

Here are some facts and laws about the way that the world works. And that can work in conjunction with this statistical knowledge. So to use an example, I guess we can go all the way back to Cycorp because I'm not going to be violating any client information at this point. But if we're predicting stock market stuff, we can run like rules and, you know, machine learning algorithms to say like these stocks are good, you know, good bets or something, or like they're outliers in various sorts of ways. And that can winnow down crazy amounts of data. And then we can end up with a more fundamental view of them.

And then we can say, all right, here are 10 ways that we have modeled about how stocks can be overvalued. Here's 10 ways that we have modeled about how stocks can be undervalued. That's part of like our core worldview and model. And then we can test on those particular things. And so it's like use the LLMs and use the machine learning stuff to winnow stuff down and then be more precise and reasons and sort of rule-based on the kinds of things where we're really making decisions or checking things out. Yeah, that's probably a first pass. Another way to go though, all people who are used to LLMs now understand that prompt engineering and providing context can give you vastly different answers and vastly better answers.

And a little bit of what's going on there is you are just trying to ground the LLM on certain things. You're like, take for granted that X, Y, and Z are true, and I'm interested in seeing whether A is also going to be true. And go read this scientific paper. and now give me feedback on the draft of the paper that I sent you. And that hopes to like get things focused in the right sorts of areas. But you can view that as largely what an ontology is doing as well, right? The ontology is saying, here's a bunch of stuff that we say are true.

Here's how these pieces are related to each other, right? And you can feed parts of that graph into the LLM and get better results there. And even if that's not your ending architecture, that gives you an idea about how there's this interplay between the two different types of systems.

VivekGot it. Okay. So so that's a good answer. Okay. So I have understood the idea and relevance of an ontology, right? Well yeah so now going back to what we're discussing just before the show started, which is clients getting pitched ontology, semantic layer, context layer, context graph, you know, fill in the words. left, Okay. right and center on consulting decks. And Then they get tossed from one team to another team. Because when they ask the actual question, which is how do you actually build one? Right? Which is the most difficult part of an of an actual engagement, Yeah. right? and obviously, you know, like there are domains where ontologies have all have obviously existed for the longest time ever, right?

In or not in RDF format and so on. But even so two questions, right? One, how do you build one from ground up? Right? Okay. and second, let's say once you have an ontology, right, how do you operationalize it? Right? you obviously gave a very simplified version of that with the example of context engineering and feeding LLM stuff in some in a better structure Okay. and so on. But in an actual enterprise workflow, in an actual agent, let me use the word agent. We haven't used the word agent enough number of times in this podcast. It's a sham, right? It's an absolute sham.

CaseyYeah, Do some SEO, get some more agents in there.

VivekPlease. Yes, AEO, AEO, agent, agent, Agent, Agent, Agent. I think now this will rank better and higher. Yeah. So yeah. so how does my agent actually access ontologically enriched data if that even means a thing?

CaseyYeah, good. Let's give another real simple use case for how you might use an ontology. And they'll talk, sort of reverse engineer what parts are there a little bit, and then talk about how you get them, you build that ontology from the ground up. So a really common method of integration is not even agent-centered or anything like that right now that I'm finding. I'm running into more organizations and some consulting firms are hitting me up that say their clients are asking them where they just have a ton of data and they don't even know what all data that they have. Okay. And so the problem is we hire new analysts or our existing analysts just don't know where to find the data that exists at our company.

takes six months to figure out where that stuff lives. And then once you do, you get to this database that's got, 5,000 columns in it and you don't even know what all the columns mean and stuff like that, right? So one way to use an ontology is just to say, okay, here are the kinds of things that your organization cares about, right? So let's say that you're an airplane manufacturer or something like that, right? You care about airplanes, you care about all the parts of the airplane so I can make like a taxonomy, a structure of like all the mechanical parts. And then there's also going to be other things that every company cares about, like HR concerns, so that you have different sub-organizations within your enterprise, and you have employees, and employees have different titles and stuff like that.

Okay, so I've got all of this stuff in my ontology well enough that now I can say, now let's get all the databases that we have. And I say, this database is about these kinds of things, and it contains these fields, and these fields tell me about these properties. So now my ontology has a taxonomy of the types of things we care about, like, know, engines and wings and whatnot. And it also has the sorts of properties that we might want to know, which is like tensile strength or cost or salary, all sorts of things that our organization cares about. So now I store all of that stuff into a, into an ontology, which I'll say more about what that means and how we build that in a second.

But then my use case is to write queries against that and say, I want to know what data sets that we have that are about airplane wing tensile strength. And I can go in there and it returns and says, here are the databases, here are where they live. And column XZ is the one that contains the information that you want. Right? That sort of thing is incredibly simple, but it's super valuable. Okay. All sorts of reasons. One, it just shortens analyst time, right? And two, you figure out, you're like, wait, all right, I find that we have like seven different columns that are giving us the same information purportedly.

Is that really the same information? Or are we using the same word in different ways? And like, they're not all the same sort of tensile strength. Is one of them just like derivative of another one? And so we can just kill that column and use it, or, you know, should we express dependencies between our tables? That will help us understand like how data freshness and stuff goes. Now you can think about how agents would get involved in that and sorts of things. That's fine. But like, I think just at a simple level, making an API call that writes it, making an API call that generates a SPARQL query that is maybe pre-generated and you fill in some blanks, it hits a running triple store and that returns what databases we need.

That's an architecture that works for having an ontology and can grow later, right? Cause then we know about the sorts of things the company cares about. So then when there's more specific use cases, we already have in the ontology representations about the parts. of the plane, so we want to do supply chain information or something like that. I can reuse those parts of the ontology. Okay, so when I talk about the ontology, how is this stuff realized? Typically, there are different serializations for the ontology, but most of the time I am dealing with .ttl Turtle files, and they're stored in some repository, and you start up a triple store.

Maybe that is Amazon Neptune. Maybe that is GraphDB. Maybe that is Apache Jena, whatever you like. I just started a triple store. I load those turtle files into it that's the thing that I'm running my queries against. And so your question for how do you build those turtle files, right? This is the build the ontology from the ground up. You've got a number of different options. I mean, you said we were talking about this beforehand. The first and most important step is to pay me lots of money. Okay. and then we'll solve the rest of the problems, right? But there is some truth to like, you gotta find someone who can build these sorts of things, right?

They don't just magically appear. In some sense, you might think they magically appear because you could take some third party ontologies off the shelf, right? So there are these upper ontologies or maybe you're in an industry, like especially in the biomedical field, there are a bunch of, like there's a bio portal that has a bunch of different biomedical ontologies in it. Maybe you can leverage some of those. But also if you're doing that and you don't have any ontology experience, you should be pretty careful. because, you know, in the same way that like, you just can't evaluate whether the stuff's any good or not, right?

You want to take this thing and leverage it. you know, I'm not a, so I'm not a programmer. have like very limited programming skills in the same way that it would be stupid for me to think that I'm just going to build my own application by just vibe coding and end up with something that like works great. We all know that's nonsense. I'm going to end up with a rat's nest. Like maybe it's useful enough to get me going and get some ideas. But like eventually, if I want to build a production worthy bit of code, I need to have someone who can oversee it and make sure that it satisfies the kind of constraints that we want for production worthy code.

You need the same thing for building an ontology. But how do you build that turtle file? You get someone who's qualified or you pull in some third party things, but the nuts and bolts of it. It's kind of the story that I told at the beginning for the airline company or whatever. You build taxonomies, you build the classes, you say here are the kinds of things that we care about, make sure I have a representation for all of them. You say here are the relationships that those things can have, so you build representations for those, and this has to fit in a special OWL/RDF syntax, but that's not crazy complicated.

And then you have your data that lives somewhere, and you have to figure out how can I get my data into that format, right? So you have to build some sort of ETL process. Now this is why I sort of suggested the metadata strategy. Because if you're doing metadata, the way to get your data in the format, I don't actually need to transform like my individual rows within my databases, right? The metadata is just, I need to say that this database has this name, it lives in this place, it's about these types of information, right? So then I'm not on the hook for migrating all of my data into a graph.

I just need to put my metadata in the graph and I can get some value out of it. But. Got it. There, that was, that was a lot of stuff, but does that kind of speak to what you were looking for?

VivekYes. It it does. It does. And the and the like yeah.

JoshuaSo I had a question. So one is, we need ontologists. We need people like you to consult. But who in the organization would own that asset finally?

CaseyThat's a great question. And I don't think there's a single answer to that question. And the reason I almost felt guilty about this for a while is like, I don't know what part of the company I belong to. And I've been at several different organizations. I've lived in like different houses within there. The reason that makes sense, I think, is that an ontology is part of the knowledge layer or the semantic layer, whatever you like, right? And the grand vision is usually that The ontology can store all of the information that our company has, like all of our data can be oriented within that. And so that doesn't live in any single place, right?

If we are going to put our HR data in there and we're going to put our stuff, but here are a couple of places that you can be. think it has a different flavor for it, right? So if the idea is that we want to build better schema, like data schemas that are more future-proofed and not overfit to particular applications, then I probably live with the data engineers, something like that. If we're thinking more that, you know what, we've got this LLM that's answering questions for our clients and we want to make it Mm. more user friendly so that it speaks the same language and we want to ground it in the kind of knowledge that our clients are going to have.

Then maybe I'm like more on the product side of things, right? I'm more customer facing. I'm thinking about what are the sorts of terms that customers are searching? How can I make those labels and those sorts of things a core part of the ontology so that's going to surface and make itself manifest in how clients are interacting. with it. Maybe the first use case that you have is an HR one or accounting financial one, in which case that's the part of the house that deals with it. However you slice it, an ontologist is going to want to interact with the engineers who are building the architecture that's going to be housing the ontology.

Maybe that's a simple thing where it's just running locally in an Apache Jena instance or something, and I could do it by myself. But most of the time you wanna have some engineers that they're interfacing with and you wanna have some of the product people that they're interfacing with so that they know how is the model that I'm building going to make, so that they know how the model I am building is going to make a difference But I would say I'm more often than not around programmers and engineers and probably have a slightly more... like a slightly closer relation with product managers than some of the engineers do if I had to pick like the modal case.

JoshuaThanks. There is

Viveka lot in there. There are so many questions now, Casey. I get why this would be why this could be sold to a large enterprise or a larger company where there is a long-term impact visible. There is a story, substance, impact, business ROI, investment outline. There is some champion, some buyer, somebody who actually believes in this, some buyer, somebody who actually believes in this, and they can carve out bandwidth for this as a project for like a few quarters and so on and so forth. Right? and maybe they're sort of they have some sort of a knowledge function, a knowledge management function, like a formal function and so on, right?

But most most small businesses, most mid-sized businesses, you know, for that matter, don't have these functions at all. Right? you know, their response to knowledge management is exactly how a Gen Z responds to knowledge management. it's part cringe and part like how boomer are you? Okay. Right? But then what we are also saying is you badly need a System 2. You badly need a system to actually make your agents functional and valuable and reliable and so on and so forth, right? it just sounds like a very painful transition for these for these kind of companies who do not really have an official custodian of knowledge per se, even if it's malfunctioned to start with, right? Have you had experience sort of working with you know such companies? What does that journey look like for them? How painful is it? you know, like yeah, what have you seen so far?

CaseyYeah. So for those of you out there listening, the journey with me is very smooth and wonderful, just so you know. No, it varies. huh. you There's a trap that you can get into if you think semantic layer is huge and wonderful and big and holds all of our company's information. And we need to get there so that we can do all the smartest things in the future. like ASI will be able to leverage our data and make us grow 10x or 100x or whatever. You can't do that. have to like, but it's, it's, it's like any program or initiative that you want to do.

have to say, what is a problem that I have right now that I can like solve incrementally and show value and then be able to reuse as much of that infrastructure as I build as possible over time so that it's not a, you know, it's, it's a, it's a single cost that has, you know, recurring benefits rather than I'm going to have to be rebuilding stuff over and over again. That, that you have to be in some position where that matters. Now, a lot of smaller organizations that I've talked to where they're like, ontology sounds super cool. Here's the use that I have. And it turns out that their use is, you know, they have data that always looks exactly the same and they want to do the same thing with it every time.

And they don't have any reason to think that's going to change. And you're like, don't build an ontology for that. Like what, I don't see what that's getting for you, right? You can just, you build a program that pulls that data and parses it in the same way every time. And it does everything that you want it to do. So that's fine. Well, I've worked with, you know, I had a, you know, small midsize clients that was selling or like renting equipment. And they wanted to both like power the retail hierarchy, or I think, you know, like any, any commercial, you know, you open your hamburger menus and you find the sorts of thing that you want to rent or deal with.

But then they also wanted to provide like smart suggestions to people. And so was, let's build the taxonomy of products that we have, but then let's also say what those products are capable of. So if somebody comes in and says, I'm trying to repair, you know, a blacktop parking lot. And I also need to like get rid of certain amount of waste. And I need to lift people up to like repair some of the streetlights there or something, right? You can say, okay, which, which machines have the functionality of lifting people up, that sort of thing. And that was a relatively small process. You can put together just enough model that says the sorts of, expresses the kinds of things that you wanna say.

I think the value that you get is, it's like anything. you're doing your own electrical work, you can probably do your own electrical work in the garage to wire up a single light or something like that. But if that's... But if you're gonna come back and reuse that later and you wanna splice off of it, then you're more likely to get fire hazards and you're more likely to realize, I should have put a different junction box at that spot, right? The thing you get from having a professional electrician is someone who's like, okay, you're doing this, that's fine. We can solve that in this way.

But if you're going there, let's do this and this, because it's low effort now and it's gonna give you a lot more options in the future if you ever wanna do that sort of stuff. That's kind of what you get from. having ontologists help you build out some of these models, right? You know, when you talk about what's a like least painful way to do things, sometimes if you want to play around with how ontologies could help you, you can ask an LLM, build me an ontology for this data set that I feed in and then practice like, okay, how is that kind of stuff going to work?

And that can give you a sense for how you might leverage the ontology and be a relatively painless way in. And then you just have to decide like, if this is a long-term project that I'm going to do a lot of stuff, maybe I should invest a little bit more now into saying like, okay, what are the corners I might get backed into here? but yeah, I don't, it sounds more complicated than it is sometimes. think my, my recent kick here has been like, an ontology is a summary or a list of the stuff that your business cares about and how those things relate Okay.

to each other. And. you can have an ontology by writing it on a Post-it note and sticking it to your laptop screen. And that's like a first pass of like, okay, you know what? This is what I care about. And that has some benefits already. Like there's a little bit of ontology that is just like data psychology and data therapy where you just, sit down and you're like, what does my data look like? What am I actually using? What do these things mean? Am I using the same term everywhere? Or am I using different terms to mean the same thing? Okay. Right? Asking those questions can be valuable to you, even if you don't do a single other thing with the ontology.

Right? So there's a little bit, and this is where maybe, you know, philosophers, like you got the Socratic, the unexamined life is not worth living kind of thing. Like un- un- uninterrogated data is pointless. So you use an ontology to make sure that your data model fits together in the right way and you know what sort of stuff is there. And that's not crazy hard. Now, But. if you want to put everything into OWL/RDF and you want to leverage OWL reasoners to generate conclusions about named individuals and stuff like that. Sure, that gets more complicated, but you wanna have a path that gets use out of the ontology long before you get to that point. And so when you're getting value on it there, then you continue to grow it over time. Okay.

VivekGot it. Okay. I have more questions now. I have more confusion now. Thanks for Alright. this. and I don't know if I'm also on time right now, but I'll take this one last one, which is because you did speak about there are domain ontologies already available, which people can tap into, can start with, and so on. how should those who don't have ontologists on their payroll who are not your clients yet, how should they think about transitioning from using a domain ontology right, to say, okay, now I want to, you know, for lack of a better word, fine-tune this a bit more and add my own, create something proprietary of my own as well.

Okay. what does that transition look like? For somebody, or what is that, you know, that decision itself, right? how sh how would one sort of think about some of those things? Because I agree to your point, program manage the hell out of this, start small, scope it properly, think of it as an intra, you know, build one, use many ways, you know, and keep on adding to it. Absolutely, makes sense, right? Low investments, you know, track ROI, You can decide. all of that stuff. Yes, good design, right? but transitioning, or like, you know, I was using SNOMED CT until now, which is which which which was good.

But like you're asking me to build my own, like how does that integrate into it? Are there clashes? If there are clashes, who will resolve it? If I call you know KCK, but, you know, SNOMED calls KC KC, like how wha what happens then? Like more questions basically. Yeah.

CaseyYeah. So the, the first thing to say is I take some, you know, ontology off the shelf or whatever, and I want to change it. My, my metaphor for that is like, you're buying a house, right? And often you're like, I'm going to knock down that wall later where I'm going to add an extra bathroom or something like that. Right. And some projects are more, more destructive and harder to do, you know, like repainting some walls. Not so bad. Right. If I just want to add some labels to the ontology and be like, great, I like the structure of this one a lot. It's just, call it something different here at, you know, at, the AmeriCorps, but great.

Then you just add, add a label. or maybe I just want to say, okay, that has a category for diseases, but it doesn't have a category for like, you know, bone breaks or something like that. So I need to create a new class that, you know, sits alongside a sibling class and then provide definitions. And those aren't. hard, right? You take the turtle files or whatever you had the ontology stored in. Maybe you just expand the existing file. Maybe you create a new file that just has your triples in addition to the other one, however you want to store them. Not super difficult. The difficult changes, right, are if you want to totally refactor parts of that ontology, especially if, okay, we've been working with this for a while, right?

And we have APIs that are calling and writing queries against the graph and now we want to change key parts of the vocabulary of the graph How is that stuff gonna work? And that's hard, but that's not any different than any other, you know software Development stuff that you're doing right? Like if you're you're changing your payroll software Or you know when I worked at Amazon for a bit, right? like if we're changing the what the landing page looks like or something like that then That's gonna break a lot of stuff. And how do you handle that? You make regression tests, you have staging development stuff, other just like good development practice.

I would say that building your own, rolling your own ontology or thinking about the ontology, you just wanna say true stuff in your ontology, right? So if you have... If your goal from the start is, want this model to say as many true things as we believe as an organization. I know I'm not going to get it perfect. I'm going to have some false stuff. cause you know, in part, even if I perfectly capture what we think right now, certainly we have some false beliefs about the world right now. We're going to have to revise those later. Right. So you have to go in knowing that some of the stuff is going to change and that you're going to add and enrich it.

And lots of the time. The best thing to do is say something that's true in general and know that eventually we will be able to say more fine-grained and specific things later. Right? So I'm going to create a class for employee. Done. Now I'm going to maybe need to put exempt employees and stuff like that in later. and that will be a subclass of that and I can handle it, but I just need to know, okay, I've got a place where that's going to go. So I, I know that I'm not blocking myself off from doing that. And then I'll handle it when it becomes a valuable and useful thing to do. I don't know how much that speaks to the issue,

VivekYeah, no, it does. but yeah. It does. It does. It does. It does it helps a lot. I think the example you gave was beautiful. That depends on, you know, like what you want to do and how much you want to do. and the fact that it's like, you know, just any other software, migration, upgrade, updates, sort of a process. Just think through it. Yeah.

CaseyYeah. One thing that's really important here, and it's a thing that like, I feel so natural to me, but it's clearly not to a lot of others in the software world, and for good reason, given the other structures, but there's the open world versus the closed world assumption, right? So the closed world assumption says if it's not true in our data model or our schema, like we can't prove it to be true, then it's false, right? Like we assume that our data model fully encompasses all the things that exist from our company's perspective, right? And the open world model says if it's not true, like not provably true in my model, maybe I just don't know about it.

Right. And that is the way that OWL/RDF thinking goes. Like I put a bunch of slots in the ontology. I don't think I've fully represented everything in the world. Right. Like I say, there's employee, that doesn't mean that there's not more specific kinds. Right. I, I asserted that Casey's an employee and I didn't give his social security number. Does that mean he doesn't have a social security number? Of course not. It just means that I don't have it in my data set. Right. I think what. People who are used to close world assumption like schema building and models they get like no I've got to get it all in there right when I first build this model and that's not all the way to think about our ontology building You're you're you're building as much of the model as you need to answer the questions that you that you need to and as you get better at that maybe you're you're gonna get good at like Guessing ahead and maybe putting a little more detail in places that are gonna be useful later, right?

But but for now you just say that say all true things, as many true things as you can that let you use the model for the purposes that you have, and that will continue to grow and expand over time.

Joshuathink programmers can get it. I guess the analogy would be like, how would you design, let's say a table, and based on the assumptions in that design, you would say that for whatever reason, you're not going to get a social security number. But maybe another opposite assumption in a different system would have said we must get this. the social security number without which we're not going to create the record. So it's, think, down to design. And if you put it in those terms, think programmers can get that it's basically a choice. There could be an open world that you don't know what you're missing. Yeah. Everything is nullable in that sense.

CaseyYeah, there you go. from on the other side too, right? I don't want to give the wrong impression here. There are lots of times when I want to close my world. There are times when I want to say, if I don't have a social security number for my employee, throw an error, right? Like that's bad. This is my HR data. I expect to have that stuff and it's a bad sign if it's not there. And there are different ways you can enforce that. There are things like SHACL that specifically look at that validation. Or you can build your own infrastructure around it that writes SPARQL queries against your database, looks for things that are missing that you expect to be there, and then throws errors. So all that stuff is good. But it's just, it's a live question whether you want any piece of data to be viewed as exhaustive or not.

VivekPerfect. No, thanks. This is the this has been very, very interesting. I know for a fact a lot of people if I a lot of software folks who have conventionally distributed software would want to, who have conventionally distributed software would want to prompt their LLM: what does the ontology look like? If you had to create it.

CaseyYeah. And there I will just, you know, reiterate my, like, I think that's a great thing to do, do it and see what kind of stuff turns out. but don't, don't fall prey to, like, you know, if you asked it to write your program for you. Right. that it would do some good stuff and it would do some really stupid stuff. Right. So don't just assume since the output is like, that compiles and that is Turtle. Like, yeah, it can make it can make Python that compiles too. That doesn't mean that it, that it's all good. So. Run with it in the responsible ways of using generative AI and not in the, great, looks like we don't need ontologists after all.

VivekI like the responsible ways of using AI.

JoshuaWe're going to post that out here at the top. Responsible.

Vivekresponsible ways. I think you sh should be my backdrop, like responsible use of AI. Not responsible AI, but responsible use of AI. That's a good one,

CaseyYeah. We have, Casey. I've seen this on a number of emails now where there's like, this content was partially produced by AI. It's like, please, please check and use at your own discretion or whatever. It's like the surgeon general's warning on a box of cigarettes now.

VivekAll right, Casey, thank you so much for this. This is this is absolutely fantastic. Right. But before we go, I know you have an amazing YouTube channel where you bring in a lot of experts and a lot of academic you know folks from folks from academia and you go at it, right? where should your future clients and how should your future clients contact you?

CaseyAwesome. Well, thank you. Yeah. So I do a few things. I have my YouTube channel, which is Ontology Explained. And I just started a new podcast with a guy, Carl, from Internet of Bugs, called Philosophy, Programs and Prompts. You can find both of those. If you go to caseyhart.com, there are links there and you can, you can find what you want about me. You can hit me up on LinkedIn. A number of ways. If you look a little bit, you'll, you'll find your way to me. So yeah. Feel free to reach out. always enjoy having conversations about ontology stuff.

VivekAwesome, Casey. Thank you so much once again. Appreciate having you here.

CaseyYeah, fun to be here. Thanks guys.

VivekWhen a model invents a fact, we call it a hallucination. And that somehow makes it sound smarter than it actually is. This is what I love about Casey's phrasing. His point is sharper than the usual complaint. The model was not malfunctioning. It was doing exactly what it was supposed to do. Generate text. It did. The problem is when we ask a text generator for the truth and then act surprised. This is exactly also why Casey will not take the answer without the why. And an ontology is exactly how you get the why. Not some mystical enterprise megaproject. An honest, written-down list of what a business cares about and how those things actually relate to each other.

You can start with it on a Post-it. You can start with just the metadata. You mostly have all of those things already with you. If you want more of Casey, he runs the Ontology Explained channel on YouTube. And he has a brand-new podcast with Carl from Internet of Bugs called Philosophy, Programs and Prompts. I love that title. Everything is at caseyhart.com. Link is in the show notes. One thing before we go. If this show is useful to you, do rate it wherever you are listening. Spotify, Apple Podcasts, this is the entire marketing budget for us. Once again, this is ContextOps. If you are building enterprise AI in production, this is the show for you. And remember, where we landed at the end: responsible use of AI, and not responsible AI. Once again, thank you for listening.

Operationalize

Take this episode to your AI

Open this episode in your assistant with the summary, key points, and a link to the full transcript — then put the insights to work in your own context.

Opens in a new tab · shares only this episode's public transcript link

Keep listening
S1 · Ep 12Jessica Talisman

Context lives in the relationships

Jessica Talisman has been building knowledge systems since 1997, when Steven Spielberg's Shoah Foundation hired her to catalog Holocaust survivor testimony, on VHS, into two-to-six-minute segments. Twenty-five years of library science and enterprise information architecture later (Amazon, Adobe, Overstock, and the Department of Justice among them), she watches the AI industry rediscover her discipline and hand it to the marketing department. Her central claim in this episode: context is a property, not an object. It lives in the relationships between things, and the document you paste into a window carries none of them, which is why a bigger window changes nothing. She walks through the Ontology Pipeline, her iterative alternative to the big-bang ontology project: define a controlled vocabulary, test it against your LLM, earn the SKOS taxonomy, then the metadata schemas and lightweight ontologies, with a shippable artifact at every stage. Along the way: why roughly three quarters of your organization's context never made it into the database, why taxonomy is the early readiness test for whether you can operationalize an ontology at all, why this is not a data problem, and why you augment before you automate. All from a guest who named her company Contextually years ago and now cringes at the word.

Knowledge InfrastructureView episode
S1 · Ep 13Giuseppe Futia

Bring the graph to the data

A citizen in Italy asks a public chatbot when their civil-service exam is. The honest answer keeps moving: dates get corrected, sessions get cancelled, and amendments pile up across official notices. Ask a language model alone and it answers confidently, sometimes with a date it picked at random, sometimes with a session that no longer exists. Giuseppe Futia built the knowledge graph that keeps the answer current. It does the deterministic cross-document work an LLM cannot be trusted with, matching each amendment to the exam it modifies, superseding old versions, cascading cancellations, and staying auditable throughout, inside a production pipeline behind a chatbot he says serves close to a million citizens. A former La Stampa journalist with a PhD from Politecnico di Torino, and the first European guest on the show, he also opens up the regulated side of his work: healthcare data that cannot leave the country, inference that runs on premise, and patient records he is not allowed to move even inside his own infrastructure. His portable lesson is the entry point. Prove value on one well-defined task the institution already needs, then earn the right to expand.

Knowledge GraphsView episode
S1 · Ep 11Himanshu Singh

Relationships should be the product

Himanshu Singh has built knowledge graphs three times at three very different scales: a politics subgraph inside Microsoft's Satori, a zero-to-one product graph at eBay, and now Netflix's Entertainment Knowledge Graph, where he leads engineering. The line he keeps returning to is that relationships should be the product. Node count is not the measure, and a graph that duplicates what already lives in your CRM or your warehouse is mostly cost. He is notably relaxed about technology choice, pointing out that Netflix built its own real-time graph abstraction over a key-value store because no native graph database could absorb their write volume. What he is not relaxed about is data quality at the point of entry, because once bad data is in a graph and connected to everything else, undoing it is very hard.

Knowledge GraphsView episode
Go deeper

Further reading