# Loophole - Adversarial Agents To Stress Test Your Morality — Brendan Rappazzo — session 2026-07-02T20:30:00.000Z → 2026-07-02T20:50:00.000Z

_44 transcript lines · 0 slides · source: full recording_

## Transcript

Awesome. Thanks so much for that. Unfortunately, we have to close it at that, but we can now decide our winner. Show of hands in the audience. If your mind was changed and your Now on Dex's side, raise your hand. Let's go. Okay. If your mind was changed, you're now on Ian and Jeff's side. Yeah. Yeah. I kinda couldn't see with the light. What did you guys think? It's pretty close. It's pretty close. Yeah. It's impossible. The delights are so right. All right. Well, that was a good debate. Thank you so much for listening. We'll see you next time. Alright. Thank you everyone for coming. So I wanna, I'll be talking about my project, Loophole, and I'm actually a machine learning researcher at Morgan Stanley, but this has nothing to do with Morgan Stanley. This is just a open source project I've been building for fun, and to give sort of the high level flavor to start, it's really this, game you can play that's built on top of this adversarial agent framework. So you specify your morals. System, and then these two adversarial agents try to find contradictions in your morals, and lately I've been building different extensions on top. But I wanted to, you know, start with sort of the origin story and and how I came up with this this idea. And so, this really started, you know, a long time ago. I had sent my DNA into '23 andMe for, ancestry testing, and I kept hearing about, you know, more recently how DNA One agent codifies that into a legal Samples can be used, of course, to help solve crimes and all these forensics and cold cases, and I was thinking about how I had sort of opted out of of everything because, you know, I was scared of the kind of slippery slope and and how my DNA would be used. But, you know, there there are certain cases that I would be okay with, and it's sort of interesting. I was thinking, like, you know, if someone could present to me case by case, you know, we'll use your DNA to solve you know, help solve this cold. Okay, so there's murder. I can sort of say yes or no and I know where the the definition of like the and the nuance of my morals are. And, but, you know, of course like enumerating all of this case by case is really sort of cognitively prohibitive. Like, there's not a good way to do this currently. And then I was thinking sort of more zoomed out that there's a lot of analogies to sort of the legal system as a whole. So, you know, one way to think of what a legal System is in our society is really just a way that we are trying to codify our own moral beliefs, and I think sort of in a similar way, like, finding the true nuance of our morals and what the law should be is this really hard translation task, and I think often we kind of on the side of being too general because, you know, finding that nuance is really difficult, and if you try to have a perfect translation, it can lead to these kind of weird corner cases or or weird failure And I even think, kind of, like, peer to peer when we're relating to each other politically, a lot of times the disagreements are more kind of fighting over core values which we don't really disagree on, instead of exploring the really nuanced points of our morals. And, I think, you know, following that kind of broader legal example, I think in the, you know, the English common law system, that's why we sort of lean on case law so heavily because we know that Finding this this nuance and these nuanced boundaries is difficult, and so we kind of rely on smart judging to interpret and apply the law, correctly. And, you know, of course, even with, like, the Supreme Court, things can get elevated, and we can decide whether a law is valid at all. And so that was sort of, you know, the idea is, like, can you take your morals, and can you do this kind of synthetic case law generation? So, you know, This is overwhelming to do by hand, but it seems like the new generation of LLMs are finally, sort of, smart enough to do this kind of high level moral reasoning. And so, that was, sort of, the the the starting point for this game, and I just wanna take you through sort of the initial release of the game, the setup, how it works, and then also talk about some of the different branches I've been building on top of this open source project because I think it could go in some interesting directions. And so at a high level, How the game works is you in natural language, and this all happens, it's sort of like a terminal based game, you specify your morals and it could be, your general morals or maybe about a specific subject, and then there's one agent that takes those morals and drafts sort of a really rich legalese codified legal system, and then it just operates in this loop where one agent is instructed To try and find loopholes in your system, so something that is, immoral but legal, and another is, prompted to find overreach, so things that are actually moral but illegal given your system, and a judging agent looks at the your morals, the produced legal code, and these sort of synthetic case law examples, and first sees kind of an auto patch, so, like, maybe the original draft of your your legal code sort of was an imperfect translation, and there's not really a contradiction, and it can just sort of auto do this update, or Maybe it's really kind of an under specification of your morals, or some kind of contradiction in your morals, and in that case, it raises it to you as the user to sort of be the judge and make a determination. Is it still on for you? It disappeared for me. Okay. So I know that Lido, if you I hope if you're curious about the game you'll play it, it's all on GitHub, but I just wanted to show some examples, and this is a lot of text, so it's more about just showing the kind of shape of the the input and output. So this is sort of how you would provide your input, and going back to the DNA example, you might, you know, specify some number of Moral principles, and then the sort of codified legal system, again, just kind of looking at the shape, has this really, like, legalese, you know, preamble, articles, sections, really trying to be, you know, write it in precise legal language, And then these the different kind of synthetic case laws get suggested. So in this case, it's talking about this is a loophole it found where an insurance company, trained a predictive machine learning model not on your DNA. But on artifacts of the DNA, and so it's saying, you know, this is actually immoral but currently legal given your system. And, in this case, it's it found that it could do sort of this auto patching, and then you get this sort of, like, git style difference of of your original legal system and then the the difference it had to make to ensure, you know, this was consistent with your morals. And then, this is a example of overreach, and in this case it found that it couldn't do the auto patch, so it's talking about, you know, Someone submits their DNA for, genetic research, but the researcher finds they have a rare but treatable genetic disorder, but currently, your morals kind of say this shouldn't be allowed, that they could disclose this disease to the the person submitting. And so this was raised to the user, to me, to kinda make a judgment, and then similarly, when you make the judgment, you get this patched legal system. And so, you know, it's just sort of a fun game, and I posted on Twitter. And shared it open source on GitHub, and for me at least it was by far the most viral post I've had, and it sort of made me think, like I think a lot of people just said it was sort of fun. You could stress test your morals, see if you have any interesting contradictions, but it also made me think, you know, is there maybe something more here? Like, could this be, you know, have more, like, practical or bigger scope implications? And so I'll just talk about three different branches on kind of exploring. The first in sort of leaving the legal area and really more practical is thinking about sort of an auto way to make constitutions for chat bots or really, you know, for agents in general where, you know, say you're a company and you want to have a, agent or chat bot that's that's customer facing and you wanted to sort of adhere to a moral code but also have things it will and will not talk about. I've kind of, in one branch, formulated it. So, you, in a similar way, write your morals, you write what the chatbot should and not talk about, and then it tries to write this codified system prompt, and then you have these kind of adversarial agents trying to get it to either talk about something it shouldn't or refuse to talk about something it should. And I see it as this sort of analogy or analogous method to GEPA, but really aimed at kind of building these codified system prompts. The, second use case That I'm I'm particularly interested in is thinking of it as a way to sort of do more ad hoc or decentralized contracts. So, I think in in a simple case, say, like, you can specify how you want your data or privacy to be handled online and you can go through this sort of adversarial game to get this codified legal system of how you want your your data handled, and if you go to, you know, say Apple releases a new terms of service or something, you can run the Contradictions between your legal system and between Apple's terms of service, and, like, surface any interesting contradictions or, like, synthetic cases where this would lead to a difference between how, you know, your morals, what you want, and what the company is doing. And, you know, in the case that it's a big company, maybe you can't really change anything, it's not a negotiation, but you can at least be. Sort of have better information about the contract you're signing. But I also think in the case of, you know, thinking more decentralized, like, if you're trying to have contracts without, you know, some central authority kind of enforcing them, and you're trying to maybe do contracts across different countries, thinking about, like, if you can specify your morals and how you want to, like, interface, you know, maybe it's just like contracted work, how you want your work to be paid for, and and and the different morals surrounding that, and the other Party can do the same, and then you both get this kind of stress tested codified contract, and then you can kind of find the the disagreements if there are any and surface them before you agree, and then you can kind of be more confident in the the contract as a whole. And the last thing, and maybe the kind of more aspirational angle is thinking about smarter government or more efficient government. I think there would be a lot of From privacy issues and logistical issues, but sort of ignoring those for now and just thinking big picture. I think for voters or constituents, you know, this could be a really interesting way of you you define your morals. You have this stress tested legal code. Sort of any new bill or politician that comes out, you could kind of run your contract against theirs and surface, you know, what are the cases you would disagree or are interesting, points that are kind of Moral to you or or a contradiction. I also think, you know, relating to one another, it's like a more I think we all have a lot of nuance in the way we feel about things, and this is a way to kind of get to that nuance instead of arguing over just values which is, you know, often the values are not in contradiction. And then I think maybe a little more practically for legislators, you can imagine, if you want to propose a bill and You can have, like, a simulation of all the other legislators in a legislative body. You could sort of stress test it before submission. And so the the third branch I've been building on this project is I I tried to do this for the US Senate. And so what I did is I first had Claude go through all current US senators and look at, you know, kind of all their voting history and anything else that was public, and build their, kind of, moral system and then Ran it through the loophole process to get a codified sort of legal code, and then on this system, you can, you know, take any current bill that's being proposed or even propose your own and submit it, and you can have Claude sort of simulate how each senator would vote. And so here you can see, like, a breakdown of, some senators, which way they're leaning, and sort of their reasoning behind the vote. And I think, you know, it's sort of interesting. Just to think about, like, seeing what, you know, a proposed piece of legislation, how people would vote, but also this sort of becomes, and I think on theme of the conference, its own verifiable domain or loop, and you could think about even kind of hill climbing the bill towards, getting like a super majority or whatever you need it to pass. And so, in this case, like this this Medicare bill I was testing, you know, it found that I think it originally started at, like, a fifty fifty vote, and it found Ways to hill climb the language of the bill such that it passed with, 52 votes, and I think, you know, this is an example of it can find, like the sort of the core tenants of the bill, and it can try to find, like, run the bill against each senator's contract and find is there any way I can change the language such that I don't violate sort of the core tenets or morals of the bill, and kind of do those auto patching that way. And then it can also find, you know, kind of rank order the changes that would need to be in place to maximize votes, and you as a user can kind of choose the trade offs that way. And then the last thing I've been trying out more recently with this branch is actually looking at, you know, kind of even bigger picture, like, can this lead to an even more Of government where you have every sort of constituent in a in a state or whatever the district is sort of have their legal code, and then you could just submit any bill and actually measure sort of the agreement between, like, the the actual voters, and so for this, I took the NVIDIA has this really great data set of USA personas, and so I took 500 personas per state, and it's supposed to be sort of well representative of the state's population. Did the same process of having them, given the persona, draft their morals, draft their sort of legal Contract and then take any bill you're interested in and kind of run it against each state, and you can also, you know, measure how much people like this bill or how much it's in agreement with their morals, and then also do this hill climbing where you kind of optimize the bill for the people. And so, just to conclude, you know, at minimum, I think it's a pretty fun game. I'm biased, but it's a lot of fun to just try Out different, you know, things you care about, put in your morals, see if there's any contradictions. You know, often it will raise some really interesting questions, and then once you kind of provide that nuance, the game, you know, the the agents won't be able to find any more contradictions and you can kind of feel good that you have, like, a consistent nuance, moral system. But I am interested in, you know, exploring could this be are there kind of real applications here for some kind of, like, decentralized or better contracts, and maybe even for legislators as a way to Sort of stress test your bills and even think about how to write, better laws that are, you know, better for the people in your district or more representative of what the people in your district want. And so, this QR code is to the the senate simulator, so I encourage you, if you're interested, to play. And, the other one is to my website which has the the full GitHub to loophole, and please, you know, play with it, fork it. I'd love to Have other contributors. Thank you. Talking about a mistake that looks harmless at first giving which is basically giving an AI agent every tool access it might ever need all at once. So, basically, that approach works well A demo. It might even work with, a small number of tools, like like, say, for example, 10 tools. But once the catalog grows, the agent gets slower. It might become more expensive and less accurate as well. That is why we are calling it the 100 tool agent wrap. In the next half an hour or so, we'll show why it breaks, what the numbers look like, and how Symantec routing with just in time context can help us. Fix this problem. So a quick introduction about myself. I am Sohail Shaikh. I'm currently working as a data scientist with Prosartica. My background spans across AI, NLP, marketing, analytics, and even engineering.

## Slides
