Video: Orchestrating Agentic Workflows: From Idea to Trusted Production | Duration: 3636s | Summary: Orchestrating Agentic Workflows: From Idea to Trusted Production | Chapters: Introduction and Welcome (73.77s), Agents and Workflows (118.855s), Agent Window Interface (186.03s), Workspace Organization (277.295s), Agent Planning Approaches (374.62s), Session Orchestration (552.415s), Parallel Sessions Orchestration (720.22s), Cloud Sandboxes (1045.4s), Model Selection (1251.29s), HydraFusion & Agent Merge (1450.635s), Code Reviews (1670.82s), Automations and Workflows (2016.495s), Agent Plugins (2525.515s), Skills and Workflows (2780.865s), Plugin Marketplaces (2870.375s), Email Testing Frameworks (2982.21s), Plugin Deployment & Governance (3141.315s), Agent Factories (3267.715s), Cross-Client Features (3406.32s), Security Boundaries (3470.48s), Marketplace and Plugins (3537.165s)
Transcript for "Orchestrating Agentic Workflows: From Idea to Trusted Production":
Hello, everybody. Thanks for joining. Where's everybody from? So I do the intro thing where I say, Charlotte reach them char woah. More turn. Awesome. Personally based right now in the Bay Area. I'm in the GitHub office. You can see GitHub perfectly behind me. It must be awesome. So let me share my screen and get started. Entire screen. There we go. So, actually, it's starting again. Let me just refresh that slide deck. I am here to talk about agent orchestration, and it gave me a really broad topic to basically talk about anything, and that's what we're gonna talk about. I think the whole spectrum, how we're moving from chat to system of work. And big part is right now how we work is is kinda there's a lot of attention being spent on agents. I think initially as we work with agents, it's constant interruption just sitting down with an agent trying to babysit it to do anything. And we have these points now where it's really about you setting the scope, the multiple agents potentially doing the work for you. You occasionally jumping in to look at the answer and you review it. And that can be across multiple stages, the research, implementation, investigation. So how do we, like, retain ownership in the system and then really, improve speed as well as we run more things in parallel? So one way to do this, and I for background, I'm Harold. I work on the Versus Code team, but I also work across GitHub Copilot. But because I like I work on Versus Code, and we've been doing some amazing work in this, area to bring you kinda more out of this code focused editor view into, the rest. So usually what most people will have right now open with the Versus Code is this idea of like the typical Versus Code editor surface. So you have a lot of, chats on the right, potentially even multitasking already, work on multiple things, like, what's this model? What's the coordinator model in that project? So you can ask questions for the project, but it's still kinda you know, you're waiting, you're sitting there looking at it. What if you now want to switch to something else? Could you close this or just say no? Actually, it's close this. So is now what we did with agent window. This is the new UI. What's also behind agent window is actually a these agents now run-in a new, process. So even though I closed my editor window, I can still, at any point, go back and, again, jump to a mode that's closer to the code. I can now actually zoom out and focus on everything else that I'm looking at. Let me just hide this. Okay. There we go. I'll switch our permissions and just keep allowing. So this this new, session that I just started in the editor now carries on no matter where I, continue on. And that's great because I can go into the editor. I can stay, kinda in the code if I need to answer architecture question or really dig into a problem or just answer high level questions like in this case, what the model is. Should be excluded. I already updated it. That's great. So that's, agent window. And in this case, I actually have a few things covered. This is where I create my slides and this is actually a group I created. So I can start actually orchestrating the things I work on by project, and stay on top of everything. So I can start curating my own workspace. I can make things more compact or not, depending on how much space I have. I can move, but I also like, grouping by time so I can see can can focus kinda treating it more like an inbox and see what's most recent. So, filter down everything else. So it becomes really more of a, inbox management process kinda project coordinator and still feels like this code. It comes with all my changes as well, multi change view, and everything else in one box. So really easy to switch from high level routing and asking questions and planning to something else. So I can even ask more questions like, what are the models? So I have one answer. I can go deeper and start actually multitasking because now this session starts with the context of the prior session to answer a follow-up question. And that's great for planning. That's great for, you avoiding long running threads where you keep bouncing around between different topics. And sometimes you just wanna answer one off questions, that don't require a whole new session or require you to, like, veer off a very focused session you've been doing. So more multitasking in one place. It's about workspaces that they all carry over everything you do, and an editor comes over the changes and the conversation. So a big part is, I've worked on plan mode for over a year now in Versus Code before, and we're looking at planning and how this evolves. But planning has been always a big part. And even as you orchestrate further, planning becomes really important, especially as planning allows you to tease out the uncertainty and the ambiguities in your project or in your question or even in your thinking. So planning, even as you go, multi agent, multi session, background agent, cloud orchestration still is a really critical part and remains something really interesting. So the ask being able to ask questions is a really big help. But then I started three sessions here. Let's make them space out. So we have, like, one is a typical planning session. So this one this question started in the plan with by pointing out a problem. The agent looked at the page and came up with some ideas, proposed something, that I could just pick as recommended. So planning, ask some questions. Or even on this session, the design thinking session, I asked it to do more diverge converge. I didn't use plan mode. I just asked it to ask me many questions. So, again, similar framing, but now it it will probably go on and keep asking questions until we both say, let's just do it. We we figured out the problem. We figured out the right solution. This one had, like, three new updates bar. That sounds good. So that's one solution I came up with, and it probably will ask more questions. This one is now actually already writing a plan. But, there's actually another one where I ask it. I gave it the same problem, but I didn't wanna ask questions. I just it was a pretty well framed problem, and I just wanted it to actually, open up the concepts. So in this case, it has it came up with three prototypes. Let me make this full screen. So three prototypes, it just and that's, I think, the the magic of working with agents. Originally, we said, oh, the new coding language is it is English, then a new coding language itself became mostly markdown and giant markdown documents. I feel now the the best coding language with an agent is basically whatever you need to answer the next big question. In this case, it's I wanna understand how in my inbox that I'm building, how can I properly, avoid kinda layout jumping around as new items come in? So now I have one idea for stage rollouts. I can even build a little stress test for me as things are coming in, and I can just try it out. So that is so much easier than me potentially answering tons of questions like design thinking is still doing, where it's like keep, preview, feed, freelance. So, click around and bring that into meeting as well. So imagine how, cool it will be as we go into a meeting. I'll share the prototype, and people just click around. So that becomes really important. And then, do we have a plan already, actually? The plan is to ask some questions. So the plan now also shows up here, and now we can actually go in and, start commenting on it or start editing here as well. So in this case, let's rethink the problem. It's mainly about attention. Okay. I get feedback, and I submit it. So that's how I wanna work on plan. I don't wanna edit markdown. Actually, most of the time, I just wanna give the agent quick feedback. Especially on a plan or a long plan, it's much easier to just add comments and keep going. K. So the planning, we talked about how to make it reviewable. It's really important that you figure out what the best way is. Let me just go back here. Okay. We're all answering questions. Those behind the scenes, because I only have one screen. Payroll work. Let's talk about more about payroll work, how we actually wanna, start more sessions. So we have session orchestration and steering. So I already showed how you can start multiple sessions. I have, like, three planning sessions that did work, that asked questions, that become somewhat of an inbox that can easily jump between and work and context switch. It's easy to keep the plan and focus as well as I work on this. So as the agent starts, getting into details, I can keep, the context in place. So that's already one place to do my work, and I can do that even across multiple places as well. I can do this in my Versus code docs and start coordinating as well. So just switching around, being able to do that. But oftentimes, it's actually about how can how can I hand off some of that orchestration? How can I not, babysit every agent and ask questions and answer questions? So in the case of my mock up, they all look kinda cool. But it can actually now say, implement, start three sessions to implement all solutions, all proposals, or do just do that fits you all. So now I have an agent that understands the problem that already explored it and that can now kick off three more sessions, in there are not sub agents. So sub agents are running within one session and run-in the background, whereas the, sessions that I'm not not creating well, if it helped, demo to demo gods. Perfect. Perfect. So that's going. So one one orchestrator, three sessions spawned, all working in isolated work trees. They don't will not fight each other. They will just work. And I could follow all of these or just follow the main one as it does its work. So it becomes a lot easier to stay on top, what the agent's doing. So we all look at design brief. They're I wanna go I wanna do all awesome things. So we can look at that later. So the cool thing, these subsessions can now actually talk back to the parent session. So I actually have an orchestrator that can make decisions for me and in the end, summarize all the work back to me. I can follow-up and have it compare and contrast and do all kind of other follow ups and just keep driving. And then, eventually, I will probably just pick one solution and ship that into PR, or I pick, the best of two solution and stack them into PR. So that becomes really interesting as you think about how do you pair run more things in parallel and that one agent orchestrate. You can have one plan and then spin off more sessions as you work on the details and keep stacking PRs. So what's interesting now actually that these are now running work trees. So once these work trees will try to run a project, they might run into actually port, conflicts and other things. That's the downside. The work trees will isolate the changes so I don't get conflicts as the agents edit in the same place. They can edit in the same area, implement similar features without running into each other. But once they start running and testing, there might still be problems. So how do we want it? Work piece don't because they they shared everything as shared. The solution would be one of them, dev containers. So dev containers, hopefully, a few people use it already. If you do post in the chat, are a great way to isolate some of that into, into systems that just run stuff containers are basically defined. In this case, for Planwork, I have one. Define what the environment for the project is. So they're checked into the repo, and for PlanWork, I can just go in, use dev container. And again, again, viewing this expression, which model is the or, maybe update model, switch coordinate model to that's with Terra. You know, it's LUNA, it knows Terra. So we picked a dev container, which means, it will run-in a container. Dev container will inform what the the environment actually is, and we just kick this off as well. So we get a dev container that we'll just launch to other things, locks it noisy. Don't look at that. But now this container can actually run fully isolated in a container, so it's much easier. It still runs on my machine, though, while the other thing so it will take up memory, and I shouldn't run too many of those. Otherwise, this this will right now might be cut short. So next up, it's cool because, like, this definitely actually runs in Kinner, but Versus Code just works. I don't have to care about containers. I don't have to care about, what what happens in these containers, how how the how Copilot is set up, how this code is set up. It all just works. So I'm trying to also, if you're busy in my notes. So, I'm just here, post a chat. Like, how many of you are already working on remote environments? So I was supposed to touch. And what I do, would be, would be at the oh, don't fail the demo gods. I I use insiders. So if you've been or wanna give more feedback for us, it is a great way. But what I've been doing actually to go back to demo is I have this machine connected. Big box is actually a Windows, a virtual machine in the Cloud that I have connected to via, DevTunnels. So dev tunnels, an existing system in Versus Code allows you to connect, to any machine that Versus Code is installed. You just put that machine into remote, and you log in with GitHub or Microsoft. And now this machine shows up anytime even now I can create a new session. So, shows up in my remote sessions, big box. I can just launch a session here at any point or go into all the sessions that have been running. So this is this one I just kicked off before, asking about how we put models in the Versus Code repo. Versus Code is you can contribute to it from anywhere, but once I start contributing at scale, running it in a DevBox makes it a lot easier. So and that's a great way to, like, move workloads off your machine, care less about the space and the memory that they take. So that's something you can always set up. So if you have a Sage machine or any machine sitting at home that you can use for development, you can already start moving your agents off your machine. So cloud sandboxes, that's where we where we see more traction in the future as we're rolling it out. It's already in preview now, at least internally. Some some people have access to it already. It's in a Git Versus code, GitHub code by the app, and the CLI. So you can start moving these workloads off your machine. The difference to, for example, the the code spaces is that this actually runs, fully in a micro VM. So it's much faster to spin up within, just basically in blink of an eye if you have an environment set up that your agent can use to either work on a project, check out a repo, or do all kind of other work. So moving things off your machine in in on demand, compute, that's the goal. And you don't have to care about it. GitHub will handle everything for you. So in Versus Code, we have a early preview of that with, that we are just shipped in insiders for internally as a team. So for us, we would switch to cloud, and then we would check the sandbox. So our very early UI, but once I kick this off, we'll see just one off my machine and but show up as everything else in in my sessions. So it's really cool to get to get cloud compute without you caring about cloud compute just as as I didn't care about the container before. So next up, we have the agent host protocol and because the question is how how would the Dolby sessions connect? Is it just some some connection via scope managers? Are we just dev tunnels is one way. But dev tunnels just allows you to connect to that machine. It's similar to, like, SSH and just making that, easier to use. So but in this case, actually, we we we are and I showed this in the first time. I could actually close my editor, and these sessions continue to run. So agent host protocol makes these agents multi client aware, and I can at any time connect my my clients from the host where the agent actually runs. So the agent window, I can even open this on vscode.dev and have it connect there. And agent host protocol manages the handoff and keeping everything in sync, including chats, the terminal, and all the changes that happen as well. So this is, it's actually an old protocol if you look on GitHub. We already put it up on Microsoft and Microsoft, and we're hoping that that's basically gonna be the underlying infrastructure of allowing more for GitHub. But then, also, if you wanna spin up your own systems, like the project I'm showing with plan of work, exit actually using agent host protocol under the hood to provide, like, an alternative orchestration interface on top of the agents that I have running. So it's really fun to, like, build on top of it and build your own agentic infrastructure potentially on prem as well. So we're in, talk about like, what people think, actually, what dev containers provide is also sandboxing, or how can you even control what, agents are able to do? So containerization is a great first step, but oftentimes, containers still have a lot of access to the host system and are not actually, full security boundary unless you configure them as such. So local sandboxing is one one solution here as well that provides the the, OS level sandboxing for agents. So local sandboxing in Versus Code and Copette CLI in the GitHub app. It's in preview right now, but basically allows you to constrain what files can be accessed and potentially constrain the network as well. Can the agent even talk to the outside network or just on the machine? And a lot of these controls in the past lift in the harness, so we would try to block when the agent access the URL by it using the fetch tool. But once it saw that, it could still easily fall back to the terminal and do something similar. So a local sandbox is one solution that you can start, isolating more of that, for security as well. There you go. The other thing we see with agents is the either you go full, auto approval or you'll as many people call it, or you approve and you keep remembering, hopefully, as you build up, that muscle of, like, what the agent can do or can do it. I think we all somewhat lie to each other. We still read every single tool to pull. So what we want to offer in this case is a much better way that you can actually, take this with confidence and take both put some distrust back into the agent. So in this new picker, we have now the assisted permissions. So manual would be I with everything assisted now means that the agent actually will start looking at the tool calls, at the context around the tool call, and assess risk for you. And then when it finds the risk is low, it will approve. Otherwise, it will ask you for approval, and then you can decide. So it's a great way to derisk, tools as as you just run and and have a more autonomous experience in these agents as well, especially as they run off your machine and you don't wanna babysit them as much. Next, I think the biggest problem people see is, like, which model to pick. And I think I I definitely have very, very many conversations with people on their preferred model at the time, and then they have preferred models for each task. If you have if you have one of those, like, what is your current preferred model and why? Appears we wanna post them to chat. But the idea is basically that most developers I talk to are overwhelmed and confused, and it's hard to keep up. So HydraFusion is meant to solve it. Auto model in this code, which we only have for a while, is routing based. So you give it a task, the agent will look at the task, and then pick a model based on the different model it has access to. So we already have now in the model picker, we have the auto tiers. That's actually I think it's not shipped yet, but you see the last new previews here. So we have the balanced efficient, balanced, and intelligence, and they will take different tiers of models, already. So balance is probably what behaves more closer to auto in the past, but intelligence allows you to I know this needs more reasoning. I still don't care about the model and keeping up with it now so or if it's, like, Luna Max or something else. So you can just pick intelligence, and we'll pick for you. So that's auto, but HydraFusion actually adds a layer on top of it. What the problem of auto, it still mostly runs on the first prompt, and then it could change later on, but it's it's, still more less dynamic than what you would want, especially as you might do first some research and then go deeper into another problem area that might take more reasoning. High diffusion allows it by spinning up multiple sub agents, multiple agents that will spin up with different models that will then delegate task to depending on what the problem is. And it's it's a really clear signal in any research that and giving them different tasks unlocks a lot of potential. In this case, we see on various benchmarks, we see for general bench, higher task quality. DeepSree is slightly lower, but at significantly lower cost. So you can really use the best of the models you have access to with GitHub Copilot many more models and get higher quality out of your results without starting to constantly switch models because you do research and then switching over because you do implementation. So HydraFusion will do that for you, automatically. Also something that works is advanced copilot or autopilot. So if you if you are in the agent window, you can always switch from interactive to autopilot, which means we will actually tell the agent to only stop when it's done. Advanced autopilot will improve that further by having another agent review the work and assess if the work is actually complete. So kinda improving on the on the promise that the agent is actually fully done and has achieved the goal that he gave it. So another thing we don't you don't have to worry less about telling the agent to continue and answer questions and keep going. So a lot of work right now was that you did all that shaping and building locally, and then you moved into to PR and then required checks run. And as if soon as you speed up that first part, that second part where you have to suddenly collaborate and get feedback can can take a toll and become slower. That's where we have investments in agent merge and code review. So one thing is agent merge. If you haven't seen it, so I probably have a few things ready now in plan work. I could show it. So this is now a work tree, where it implemented the concept a, which is kinda cool. So I could create a start to create a PR for that. So I can click here. The agent will fill in all the details, and then I click agent merge here. And I can have full control about what agent merge is able to do. So it can even start merging my my pull request. It can start addressing reviews if I feel comfortable with that depending on the project and my risk posture, fix the eye flares and resolve merge conflicts. I think the latter one I'm most excited about because as we move faster in these projects, merge conflicts pile up because people are learning a lot more code faster. So just merge conflict resolution is probably my favorite feature here. So the in this event, the agent merge will eventually also land on github.com financially. So now it can you can just have a PR and depending on how much effort you spend beforehand to to build feel comfortable with that PR because you did a lot of planning, look at prototypes, guiding the agents to the best outcome, and then the PR is just a final kind of quality check. And maybe you pull in more humans to do the review, but, eventually, the agent merge can help you once that review is addressed, to just get everything in. So the other recent improvements has been code reviews. They now allow light and balanced efforts. So in the past, code reviews just ran. They always were great because they already used, like, a multimodal system, but it was hard to kinda dial in if you have a more, like, internal project where reviews are maybe important just to keep quality high, but they're not as critical versus, higher, risk project that that requires a more balanced approach or higher, cost for reviews. So now you have a way to dial it in within your reviews and override it per repository even. So it's a great way to I think many companies I talk to are super excited to roll out, agent reviews because once developers start gaining more velocity, those reviews really help to build confidence and what's shipping and unlock that velocity. If things get stuck in the PR queue, you lose a lot of that initial velocity, and CodeWeView can help with that. The LycoView, as I mentioned, has a multi agent system. So you actually get the combined review out of this from from multiple agents that will share the perspective, and we see the quality of comments really increasing through that. So there's a really cool experiment block if you wanna read more about that. So now QuickView can actually, test this finding, which is a recent improvement. So because it's an agent, it can now actually run its code, and any improvements that it does apply or suggests it can actually have higher confidence because it has has now access to to running these problems itself and rerunning and trying. So you have a much more proactive code review similar to how a reader would actually run the project and then give you feedback as well. And lastly, something on the Versus Code side we sort of accept because Versus Code is on GitHub. So we've been investing. We adopted code review early on and took our time to really get comfortable with it, but now at some point we many months ago, we switched over to code reviews are now mandatory, and, a human will only get pinged about your code review, when all the code review comments from the agent are addressed. So bringing in your skills in MCP is a massive boon to ground code reviews and how your developers work as well and, potentially, also how your organizations work if you have these shared skills and shared plug ins across the organization. So it's really, important to build out those those skills, those instructions in your rebuild about best practices and then bring in MCPs to bring in any external context like security audits and any other services you need for to do good code review. And lastly, as I mentioned already, you can get approvals in as well. So if if you have projects where things are more, more fast and loose depending on if it's an internal project where you just built internal, tools, then you can start actually after code review is found no issues. You can start bringing those into a project as well. Okay. I'm gonna pause here quickly, unshare for questions so so we don't get them all on the end as we go over agenda workflows. Perfect. So currently, hydrophied, hydrophied, there is a blog post. Yes. Sounds good to have, so we have some good answers here already. So hydrophied is definitely I recommend the blog post, and the CLI already has it in its experimental flag. So probably the CLI right now is the best place to try it out. It's shipping right now also in the GitHub Copilot app and use code soon. We're gonna talk more about skills and everything else as well, so stay tuned on that. Okay. That was my break too. How do you drink this? K. More questions. Can you get more information about bitrate on a plan design thinking? I mean, all all the tools, I'm gonna get a code by that, this code have, built in browsers and built in planning tools. So, like, one of my first thing is, like, the idea of having, a, skills to handle that. That's that's one way. But, b, just, agents become so much better at just being aware of how you prompt. If you just ask a question, they will not necessarily take action. But if you tell them to do something, they will try to do something. So, it's it's an interesting one for, doing that. But if if you want, there's a few and and the repo I'm gonna share later as well. Now this code team could repo. But overall, assume that the agent has an ask questions tool. The agent has access to a browser, and you can get creative on what to do with that. It could be prototypes. It could be data visuals. If you just ask the agent to visualize something, with a higher reasoning model, it's usually now doing an amazing job and will blow you away. So that's kinda my my main lesson from that. But, otherwise, basically, assuming that you don't let us just need to use the agent for coding, but also to create more artifacts that you need to go along the way, that's really, that's what's wanting to think about it. Okay. I'll take over screen sharing again. I cannot take over it there, bro. I'm also also wanna check because I don't see the chat at all because I'm on one screen. Make sure everything works. There you go. We're back. Okay. Let's talk about automations. Like, we talked about lots of stuff happens in in the agents window, or you still prompted or you orchestrated it and then it runs. But Versus Code actually has now automations built in as well. So if we go back to my agents window, I have my automations up here, and one of them is a weekly developer signal previewing. So it's using I think it's actually MCP in this case to look at different areas of, like, what people talk about because so much is happening every week and I, as a PM on the AI space, just need to be able to to figure out, like, what what's happening and why why is it happening. So there's a whole history of past research where the super, interesting insight of the harness became the product surface because, yay, we have alignment of h and some d. Very excited to see that one. And and all the other parts, there's a good one as well here, skills over MCP was in here as well. So MCP skills, which I've been following as well. So that allows me to stay up to date. I don't wanna prompt the agent on a weekly basis. This cap automations will do that for me, and I can read it in my agent's window as an inbox. So that's one way to do automation. So these still run, you can pick a repository. You can run them outside the repository just doing things for you. It can reach into MCPs and skills. You can tell tell tell to not modify anything and run it on a schedule and see the full history. The other way we're working on is bringing those into the cloud. So I'm trying to talk about cloud sandbox for a bit or running them even on on my dev box that I mentioned. So my dev box actually also runs automations, and they show up as well within my agents window. Because my dev box always runs, it doesn't hibernate, it doesn't close. I can actually keep that connection up and just get all results from that automation back into my agents window. So think about if you have the system sort of running an SSH machine that is always on, you can move your automations to those systems as well. So but, eventually, I think most automations we see are moving off your machine because sometimes you wanna schedule them. You don't wanna, you might not have a machine that you wanna set up just for a purpose. Like, I have a dev box, but not everybody has, and you wanna trigger them potentially from a repo or something else. So you wanna start building these, like, autonomous, automations. And that's something we already have with the Copilot Cloud agent, and that can run inside your repository, and just work. So that's something I'm working towards is that the sandbox I showed before becomes your place for running automations as well, including the same systems and they just transfer. But that could be done, still mostly for you, and that's a very so you see a lot of exciting things. So as you try out automations, like, think about what in your day to day can you automate a way and just have the agent do for you on a schedule on a daily basis. And that could be issue triage. That could be checking your inbox, that could be keeping your, rebuilds up to date that you have locally checked out. Those are kind of my use cases. Okay. Agentic workflows as a new topic has is super exciting for me because that that's a project by GitHub next that's got a lot of traction, and we're actually actively using it within Versus code as well. So one of the demo repos I have is actually not a demo, but it's a repo I I'm maintaining. So this is h n r c. It's an open source project to get your repo ready for AI. It's been started. Basically, when we talk to customers, they struggled about how do we like, skills and instructions and agents.md. How do we create those? How do we know they're good? How do we know creating them makes things better versus worse? And how do we roll these out across a large team? So agency is trying to bundle that, and I think, eventually, all of these should be available within GitHub as a feature, but allows us to experiment and give people solution now and give them something they can deploy in CICD and or one of one repo or scale as well. So So that's a project. I'm I'm one of the retainers. It's not my main project, so I really want to make sure we we can run this continuously. So what we have set up here is workflows with, generic workflows that now I have this readiness report running every Friday. It has a deadline as well. So after six months, I either go in here and revive it or we'll stop so I don't have these runaway automations that would just consume tokens. It also has max turns and an AI limit, so it doesn't burn my budget. It actually runs on the budget of, like, the the what's the repo's assigned to, and it has very tight controls about what it can do, which tools it has access to, which are safe to use, what are safe outputs. So you see in my automation in this code, the agent had a lot more freedom with, these automations that can now run-in the cloud, much more autonomous and automated. Or trying to lock it down a lot more about how they run. This case, the agent creates readiness support, every Friday and then post this as an issue. And I might show it after I get to my login. There you go. So we see the also, if it fails, I was actually playing with that earlier. It's really nice that I I actually see an error as well. So if each agenda workflows fail, it's because it's all based on actions. There's a lot more tooling around. If something fails, you can start debugging it and improving it. So also a lesson, as you said, of these automations is how you treat them, start treating them as as engineering primitives, as pipelines as well. So the other one I have is set up as issue triage. Again, I will look at issues and and run actions tend to to triage them for me. And the good part is all in Git. So you create these in this very human readable markdown format, and then they get compiled down into these workflows and locked so they don't change over time as well. Let's see. Example. So we have a bunch of stuff we can run. So we have these code animations. Everything is saved locally. It's in your session. It's it's for you. Like, that's a place where I would start automating things away from your plate. So the the host has to be available. So if you wanna move it off your machine, you have enough options with dev boxes and remote SSH right now where you might be already running things remotely. And then lastly, we have, the middle of cloud automations, which are coming along with the cloud sandbox. We can start moving, automations into cloud sandbox. It's already available in the GitHub Copilot app, and you will find it across all co Copilot surface systems eventually. So it's it can now start to actually make changes in the things there as well and create PRs. Then agentic workflows where you wanna invest as a team, these are things you wanna automate in the repo. Oftentimes, these are processes or chores that right now take time from people. They may be taking energies because they're they're complicated or just take, take effort and take away from the the parts where you actually create value, and enjoy your work. So that's where the team actually for Versus Code, we automated our issue triage. We automated our errors telemetry, as well and set a lot of parts we keep automating, like maintaining our dashboards, maintaining our planning backlog. So all of these are becoming more and more automated agents. So we, in the end, are more the readers of, and the arbiters of taste, how they best work. And because they're versioned, they really become agentic infrastructure for us. So my favorite topic, agent plugins. One of my favorite topic, so I, I was early on in in MCP, kind of part of the early steering committee, no longer doing that because, there's a lot more work happening and many more extensions happening in MCP, SIT, skills for MCP that just merged. Still following along, still, closely connected, but I'm also a core contributor to the agent plug in spec, and I've been working on skills and custom agents and MCPs and ops and such. So the clarification I always like to make, so skills, I think, is a lot of people that a lot of people see that people, like, potentially install from and get shared from different places. And plugins actually are the packaging mechanism. But to kinda bring it into into perspective, so instructions is like the agents and the files that you have in your repo. They're meant to really become these project conventions and oftentimes in the repo. And then skills become these workflows. So we have, for example, one, for for issue triage that defines how we best triage issues. So they they capture all these, specifics of how we want to triage. Like, what do we care about? MSP servers are all about the context, tools, and data. So MSP servers, I think in many companies I see is bringing the business context, bringing the developer services into your development workflows. Some of these might have CLIs, so they could be just a skill, but many of these require authentication and require more setup, so an MCP makes more sense. Or they are. Yeah. Things like Sentry and GitHub as well. So you can use the GitHub MCP or you can use the GitHub CLI depending on your environment you run-in. Customizations are I was surprised that they still have heavy growth. So custom meetings are great if you wanna reduce down a workflow to a specific role, specific tools and model. So skills are often invoked by the agent on demand, so they they have to describe when they need to be used. But custom agents, the user can actually toggle into. So they become way stronger, or more deterministic than skills in some some cases. So definitely very similar and very, you can replace probably one with the other, and you can do what custom agents do with a skill. But custom agents give you kinda more control, and you see custom agents, for example, in the drop down here. So these are, this is my chief of staff for Versus Code that does more of an orchestration work. So when I toggle into this one, I don't have to tell it to do orchestration. We'll just, like, create more sessions and monitor them for me and report back more like an inbox. So you can do a different patterns that are could be in skills, like, in custom agents that then are actually pretty sticky because it's this skill will not this instruction will not be hooked up in over time as I change away. And lastly, hooks, they're scripts that run as per life cycles like pretool and posttool and allow you to, either apply governance, observability, data transformations, lots of different use cases for hooks depending on what you aim to do. So all of these are agent customizations and, with different outcomes and different, purposes. But they all come together in agent plugins. Agent plugins as a one point o spec came out a few weeks ago and tries to define kinda this slightly cluttered plug in ecosystem that's already out there. In its initial state, it supports the kind of definition and plug in adjacent, and we have skills and MCP. Everything else like hooks and custom agents goes into the the namespace folder for GitHub Copilot. Automation is actually it's in here because it's both recognized in the top level folder and in the namespace folder, but I would probably right now recommend it to put it into a namespace folder until we have it in the spec itself. A lot more things happening in the spec. If you want to check it out, the agent plug in spec is open on github.com. Lots of discussions and planning happening there. Right now, we have some open PRs on getting icons and more metadata in. There's also work on hooks as well. So trying to close the gaps to what's in the ecosystem. Really important about hooks, it's about workflows. So if you think, for example, Figma doesn't necessarily publish its, its MCP standalone anymore, but they publish Figma as a plug in where the MCP is in the plug in, and then the skills define how the workflows that you can do on Figma. Because Figma has multiple products out there including FigJam out of there. So the the skills basically infuse the agent with the latest best practices and workflows for these platforms. So the the agent can be effective in using Figma. Otherwise, it just relies on, like, what tools are available on Figma, and then it relies on its probably training data of what it knows about Figma, which as your product evolves, can heavily change. Also, maybe Figma eventually decides that MCP isn't the right way, and they maybe also play the CLI. So that there's an MCP in there. It's just an implementation detail. So, like, Vercel, for example, is betting on CLI. So their Vercel plug in is mostly skills and that refer back to the Vercel CLI to do all these actions. So they really abstract away a lot of the technology and stack choices that developers make as they build these systems. So one one thing is I think important to highlight is that I see a lot of teams where a few people have a bunch of skills that are really, really amazing, and then they're hard to get around to everybody else or they they drop them into folders and drop them into, shared systems. But the way to do is actually at enterprise scale, you wanna build these catalogs, these marketplaces. So you wanna treat these systems as an actual engineering system. So you need to maintain your plugins. You need to make sure you have reviews built in, you you maintain quality, and you maintain trust and to actually do the thing they are supposed to do. So then you can put them into a shared marketplace, a plug in marketplace, network marketplace adjacent into repo, and people can install them into their clients. One of the examples let's put that here. One of the examples is Versus Code TeamKit. So that's an open source repo. You can check it out under Microsoft slash Versus Code TeamKit on github.com. And it's it's kind of things that we like across the team that we found useful. So one thing that we I learned early on is model council, which, takes multiple models, into sub agents that will analyze a problem and then come back if they agree or not agree, and then even you can ask them to also also debate. It both works for planning. So they will look at a plan and propose different plans, and then they will argue about which plan is the best using different alternatives. So it's a great way to unlock that kinda hydro fusion idea, but on a specific task, and you have more control over it. I also have we also have actually emails in there. So I'm just right now actually moving from Wazuh to Valley, both are open source projects by Microsoft to do emails on skills. So if I scroll all the way down, I was picturing on a bit. I would hope to have it merged for a demo. Oh, I could actually merge it. Let's merge it here. Ready to merge. See? Live merge. It's the best demo. So now we're on Valley. Boom. Okay. We're on Valley now. What does it mean though? So, actually, Valley runs evals on each of these skills. So I run, review areas actually has one that's broken. So those are something I worked on, And it tells you how many tokens these evolves took, how many its duration, but it also tells you 40 skills did get invoked at the right time. Because if you define a skill, one big part is the description, does it get invoked at the right time? And then, of course, the other part, once invoked, does the skill actually do the thing it's supposed to do? And that's what these emails are meant to look at. So if you have have questions about emails, Valley is something you should check out. Wazuh is the the original version of that. So but the key takeaway there is still, if you dis if you distribute skills and plug ins and any agenda customizations to people, make sure you treat them as engineering primitives. So emails are really important. So in this case, in my case, I had it. I actually have this from where the RubyAerial skill was invoked too often, which means every time RubyAerial is is loaded, that's a bunch of tokens that end up in the agent where then sees, oh, it is actually not what I needed because it's not what the description says. So people really hack or rub around with the description so it's loaded as often as possible, potentially, and you wanna avoid that, but make sure that cases where it shouldn't be loaded, the agent doesn't need to look at it. So that's what emails need to do, and that's what you should do with emails. Emails are still hard. I'm not saying you should do go full in on emails and expect all your quality problems will be solved. But they're a good smoke test to ensure that your description, does what it says in at least in the situations that you can come up with. And as people report problems, you can bring those into your emails to actually hill climb against the emails that people reported on. So it's really like, every like, everybody working on AI, that's, you need to have your offline emails and probably also have metrics. Okay. So how do you roll this out? Now we have this amazing marketplace, like our team to marketplace. And now in manage settings in GitHub Copilot, you can actually roll these out across an org, even across, several teams. So within my settings, it's a dot github dash private repo, and you can point your, basically, all your clients to a known marketplace. In this case, it lives on GitHub, and it will auto update because it's the best plugins and they will evolve over time. I can start actually enabling plugins for everybody in the system as well. Any plugins will be force enabled and force installed for everybody. So that's the governance layer. So if you think, like, strong engineering primitives that every agent needs to be aware of at all times, then that's the place where you wanna enable it. If you have any thoughts and issues about, managed settings and what you wanna see, maybe more nuance, please post them in the chat, and I'd love to see them. That's the system I'm also working on that I'm very excited about, how we can evolve it further and make it better. So as I mentioned, the the interesting idea is that yeah. I already mentioned that. The if you have a skill that does interesting things that needs access to data, you can actually ship the MCP with that as well. So that's something important shift. Like, a skill is one workflow, like a a more atomic system. And I would still if you have one skill you wanna get to everybody, I would still package it in a plug in even though it has no other skills or MCPs in there. But plugins allows you really how to have that user facing description, the the distribution that you can update skills to have a version number. They don't have a description for humans. They don't necessarily have a name for it that is human readable as well. That's coming all with plugins and plug in marketplaces. So don't build skill registry. If you want to distribute these things, build an MCP marketplace. Have a few insights on agent factories. Agent factories are heavily in preview. There's a lot of things happening. The Copilot SDK has them in experimental, and, basically, it allows you to build these systems that define multiple stages and verification steps in a pipeline. So, agent could, for example, have a really agent verify agent that that then work together across, a whole workflow. If you wanna learn more about that, check out the Copilot SDK, or Copilot CLI has slash factories now as well. Maybe it's only staff shipped, but it's, coming along fast. So really exciting to see more of these primitives that are more scripted than indeterministic like skills and hooks. So you can build this right now with hooks and skills, but we wanna give you primitives that really scale to allow you to reflect, complicated workflows where it's handing off between different stages and different agents, and that's where factories come in. So, in the CLI is slash factories, and then you can see which of these different phases is running, so and which which phase it's in and which agent is currently available. What's interesting, and next slide, you'll see that they actually all keep context. They they write into context and you see what what happened in this specific agent, which is really important. Mentioned this for automations. You wanna make sure that you can actually improve the systems over time. And for agent factories that's built into the box, as they run, they will drop more context that you can later on use to improve the system, create evals, and improve them over time, which, for something that, will constantly work potentially in the background, it's a really important system to have out of the box. So the what was the end to end? We have a little moment of time. Otherwise, it's it's all about planning, planning together, creating more artifacts, splitting the work across different sessions, across the cloud, and how do you can you review and merge it faster with confidence and underneath all these packages that repeat. So we need to the takeaway is think about the the skills and plugins you wanna ship, make sure they have emails and they come together. I'll share my screen. What is the top question we have? Anyone does? Otherwise, I'll just pick something. I guess Well, all this is yes. That's a good so that's we've I mentioned I work in this code, but I should work on it across GitHub Copilot. So a lot of the, features you saw, teams are working across clients, across everywhere. So while in the past, you definitely saw some divergence between clients. We're all still exploring. We're all still trying to figure out the best UX. But as these features come to GA and actually shipping, that you would expect them that they, closely follow each other across all clients, including the one I just mentioned, is manage settings where we wanna put a private governance. So those things such as out of the box, will work reliably across all clients as well as you can expect. Yeah. So dev containers, local sandbox, like, do a lot. So local sandbox is really just tool isolation for network and file system, whereas dev containers allows you to define the environment, where you develop. So dev containers by default just take a Docker container, whatever you define, and run the agent in that. But those containers out of the box don't necessarily come with all the security constraints, like they can still access the Internet just as they want to and all that all of that. So you can actually combine eventually a dev container running in a in a or a dev container based container where it's just like this is the environment, this is node, I need, post class and all that. That was the dev container device defines, and then local sandbox on top of that could define what the agent can actually do, what files it can access, what it can read, what it, cannot write to, and all of that. So that's the security boundary. Local sandbox is about security boundary. Dev containers is about defining the development environment. Great answer by GitHub. Yeah. Don't post the spec it one. No. The spec it one. There you go. So, spec it is one of the many ways that you can define, like, an AI SDLC. And, yes, it totally makes sense. I think it's the wait. That's how it goes. So you wanna use a plug in. Yes. To your first question. Create a marketplace. Anything that's reusable, anything you wanna maintain and improve over time should be in a marketplace. So it's installed from a centralized place. So and then, it's this case, maybe that's what vetting, but also make sure it's updated. And, also, in this case, that it shows up in every client the same way. And then you can still if people still wanna install their own spec kit, they can, but you can also, as an enterprise level, say people cannot install plug ins from other places if you feel that's needed from a security level. So, hopefully, all the other questions get answered in chat, but we're on time. Hopefully, if you check out the recording, feel free to, open issues on Versus Code and GitHub. We're all out in the open source and love that. We also usually have discussions open for any releases that are coming out, so drop in there as well. Hopefully, yeah, you have a chance to try the agents window as well and get some some better ideas there how this all works and then how to run more things in parallel. Thanks for all the claps and the love, and thanks so much.