Learn to build scalable, event-driven AI agents on Kubernetes and Knative. See three agents collaborate locally, sharing context and memory, all configured with minimal YAML.
Overview
When your agent is done with its task you want him to scale back to 0 replicas and you want to wake him up with an Event Driven approach, all of this configured with 5 lines of YAML : This is what SAIL (Serverless AI Layer) is, the new Open Source framework based on Kubernetes and Knative !
During the demo we will deploy 3 agents using a super streamlined Custom Resource , kickoff the first agent with a Kafka message and see the magic happen with all the agents starting collaborating together, sharing context and memory. We wil show how all of this can run locally on your machine, leveraging Kubernetes with Docker and inferring using Docker Model Runner.
Video
Transcript
Generated 3 months ago
Summary
Generating a talk summary...
View full transcript
Okay. But, again, I'm a big fan of let me go for my through all my windows. I'm a big fan of Kubernetes, but also all the ecosystem around it. Does anyone knows about Knative? Raise your hand.
Yeah. I know your hand. No 1 else knows about Knative? Okay even more serverless. For me, Kubernetes is already serverless.
Okay? You don't deal with machine infrastructure. You say, hey. I have this workload, this container. I want to run it.
I have to create a deployment, a service, maybe a config map, maybe a gateway. Okay? With Kinetics, it's the same that you say, okay. I just define 1 resource. It's called a service, a Kinetics service, and I just point to a container, and it gets deployed on.
And the great thing is if I don't get any request, I just kill back to 0. And I wait. When I get an HTTP request, I just wake up. If I get thousands of requests, I can just kill to maybe 10, 50000 to pass, and then I just go back to 0. Okay?
So Kennedy is ready to wait to go if you want to do some service on Kubernetes. And since I already have a, let's say, a container, I just can let me go back here, and let me show you how it looks like. I installed k native, and there we go. No. That's the that's it.
Incorrect 1 again. That's There we go. K native, look at this. 11 lines of yamos. I said I want the service of the kind of k native for this docker image.
Okay? And, let's deploy that. So I already deployed that. And if we go here, and I go just I do a wedge on get pods. Okay.
I'm on the incorrect. Cubans. Cubans, in native. So ring okay. I have some thoughts here, but I don't have my my chatbot running here.
Okay? And let me do a crawl, and here I go shamelessly around my history because I'm really lazy. Okay. So here I say, tell me more about Sprint and keep attention. Oh, look at it.
I got here a new pod getting deployed, giving me the answer. So that's awesome. Okay. And and that's the funny part with Kinetik, because after 1 minute, it will go back to 0. That means that I have to talk to you during 1 minute.
I can also take a pause and drink some water. Okay? But I got my answer here, and look at this spot. This 1 here, the Doctor Mirror for Tucker, model runner, Kubernetes test. I'm really good at naming stuff.
55 seconds, I think. I put the link the the natural limit is around 1 minute, but because you give an answer, it should be a bit more. And let me just be 543. 3. 2.
1. Yay. And it's just germinating. And now my fault is just my component is at 0 plus. Okay?
All of this, just to finish, and then we can have data. That was a starting point for me. To be completely honest, I was, fired from my latest from my startup by the piece of shit, end of last year. And I have some time left, so I worked on some cool stuff. And the result of that is something called, why I I prepared all my my my my stuff here.
I got something. If you go to soltech.org, that's the ultimate result. Okay? Because I just show you here at the beginning, Knative, some stuff. But with Knative, you can wake up a workload with some Kafka messages.
And what I created is also an operator, some CRDs. And, basically, it's a bit I'm a bit competing with you guys at Docker. I'm just on my phone. But just with a few lines of YAML, you can deploy a lot of service agents. Okay?
They can communicate together. They are persistent. They have a a shared memory using Redis because memory when you scale back to 0, what's happening with the memory? Well, we are using Redis for that. So, go there, still tech.org.
It's just an open source, completely side projects that that I'm doing on my own. But I just wanted to show you the beginning of FireReflection. Hey. I love Java. I love Kubernetes.
I love Kinetic. I love agentic workflow. Let's see what we can do. And slowly, I came to this, kind of, framework. So, yeah, that's it.
And, yeah, I don't want to bore you too much longer. I don't I want you to enjoy pizza. And for those that are at the box, other piece, a slice of pizza. Thank you so much. Thank you.
You have questions?