AI vs. Financial Modeling World Champion: Who Wins, and What Can We Learn from It
In this episode of The ModSquad, Paul Barnhurst, Giles Male, and Ian Schnoor are joined by Excel and Financial Modeling World Champions Michael Jarman and Diarmuid Early to discuss a unique “man versus machine” financial modeling challenge.
The group breaks down their experience testing AI against expert modelers using real financial modeling and Excel competition cases. They explore where AI performs well, where human judgment still matters, and what the future of financial modeling may look like as AI tools continue to improve.
Expect to Learn
How AI compares against expert financial modelers in real-world challenges
Where AI adds value and where human judgment remains essential
Why reviewing and validating AI-generated models is becoming more important
How AI tools are changing the future role of financial modelers
Why strong modeling fundamentals still matter in an AI-driven world
Here are a few quotes from the episode:
“AI can build something quick, but if you want it to be right, often the majority of the time is in checking it.” - Michael Jarman
“The most interesting question is how do you get from something that works to something you actually feel confident in.” - Diarmuid Early
AI is changing the way financial models are built, reviewed, and improved. While AI can increase speed and efficiency, the best results still come from professionals who understand modeling principles, question outputs, and apply human judgement. The future belongs to modelers who can combine technology with expertise.
Follow Ian:
LinkedIn - https://www.linkedin.com/in/ianschnoor/
Follow Giles:
LinkedIn - https://www.linkedin.com/in/giles-male-30643b15/
Follow Michael Jarman:
LinkedIn - https://www.linkedin.com/in/michael-jarman-excel/
Follow Diarmuid Early:
LinkedIn - https://www.linkedin.com/in/diarmuidearly/
In today’s episode:
[00:00] – Trailer
[02:33] – Meet Michael Jarman and Diarmuid Early
[05:00] – Setting Up the AI vs Human Modeling Challenge
[10:00] – How AI Approaches Financial Modeling Tasks
[15:00] – Comparing AI and Human Modeling Approaches
[22:00] – Where AI Performs Well and Where It Struggles
[30:00] – Why Human Review and Judgment Matter
[38:00] – The Future Role of Financial Modelers
[45:00] – Final Thoughts on AI and Modeling Careers
Full Show Transcript
Host: Paul Barnhurst (01:03):
Welcome back to another episode of the Mod Squad. This week we have a special treat as we're joined by two guests, not one and two co-hosts. So it could get a little unruly at times, but hang with us. Let's start by introducing our co-host. As always, I'm Paul Bartenhurst. I host the Financial Modelers Corner. I get to co-host the ModSquad with Giles and Ian. Giles, why don't you take a minute and introduce yourself to the audience?
Co-Host 1: Giles Male (01:27):
Yeah. Hello. Hi, I'm Giles Male. I am a Microsoft MVP involved in the Excel UK Championship, and I spend most of my time playing around with AI now.
Host: Paul Barnhurst (01:39):
It's a pretty good
Co-host 2: Ian Schnoor (01:39):
Summary from what you've told me. That was brief, Charles. That was good. I'm going to try and stop you. Ian, nice to be back, Paul, and to see all of you. I haven't seen all of you in person in a little while now. And yeah, heading up the Financial Modeling Institute, FMI. Excited. Just realising that all of this group, everyone in this group is part of the Financial Modeling Global Leaders Council. And our first report came out and it's been getting some really nice traction and fanfare with people that I have been speaking to at the Big Four and banks and other places. It's being well received. I hope it didn't embarrass any of you. None of having your name on it. It has, but we I think got some really nice thoughts and we'll have to see how reasonable it was in terms of the predictions about the future of modeling in the world of AI.
(02:29):
But yep, that's me, Paul.
Host: Paul Barnhurst (02:31):
Thanks, Anne. And speaking of that, if anyone wants to listen, the conversation came out Tuesday with N&I about that report where we highlighted it. So if you don't, I encourage you to read it, at least read the executive summary, but if you're one of those like, "I'm not going to read it, but I'd love to know what's in it," go listen to the episode. Good way to get the highlights. All right, so now I want to welcome our guests. I think between the two of them, we have what, three, four world championships. We got LeBron James and Michael Jordan. Okay, so those aren't quite the names we have. We have Michael Yarman, who's a world champion in Excel. So Michael, why don't you take a minute and introduce yourself?
Guest 1: Michael Jarman (03:05):
Yeah. Hi everyone. Michael Charman. If you have not come across me before, I'm the reigning world champion of the financial modeling World Cup, previous world champion at the Microsoft Excel World Championships. I'm now a VP at a private equity freight rail fund. And yeah, very happy to be here. And the reason I am here is because, well, I had a little bit of a battle, but I suppose we'll let Dim introduce himself first. Dim, do you want to introduce yourself?
Guest 2: Dim Early (03:29):
Sure. Hi, I'm Diermer Early and I'm the opposite of Michael. So he's the reigning FMWC champion. I was the one before him and he's the former Xcel world champion and I took over from him. So we traded places last year. Yeah, I spent my career briefly in consulting at the Boston Consulting Group and then Deutsche Bank as a business manager. And I've been running my own firm doing sort of automation and data analysis since 2019. Yeah, coming up on seven years now.
Host: Paul Barnhurst (03:59):
So I'll start a little bit of the stage and then I'll let Michael and Dim share more of what they did. So basically we wanted to bring them on. They had a kind of a man versus machine contest is what I'd call it, where they worked with Shortcut, which is one of the AI tools out there and had a babble, I believe they're in San Diego, right, Michael?
Guest 1: Michael Jarman (04:17):
Yeah, yeah. Nico from Shortcut came down from SF, and we flew Dim across from New York as well and yeah, found a nice studio and did a man versus machine battle, as you say.
Host: Paul Barnhurst (04:27):
So why don't you go ahead, Michael, you can walk us through Dim, just set the stage for what you did, a little bit of what the event included, and then we'll kind of go from there.
Guest 1: Michael Jarman (04:35):
Do you want to go, Dim, as you chose the cases?
Guest 2: Dim Early (04:38):
Sure. Why
Guest 1: Michael Jarman (04:38):
Not?
Guest 2: Dim Early (04:38):
Yeah, so the idea was that the competition was based on some of the competitive cases. It was based on one round from the Financial Modeling World Cup, which is typically a short, a medium, and a long challenge, like financial modeling challenge. And then one problem from the Excel World Championships, which is a short 30-minute problem solving challenge. And yeah, I guess the idea was these are somewhat, and we can talk about the limitations of the somewhat, but somewhat close to real world financial modeling challenges, but with a sort of definitive agreed set of answers as opposed to is this a plausible range kind of thing that's more complex to assess. So it was sort of, I guess, optimised for being able to get a clear answer, as you'd imagine, man versus machine being the objective. So like Michael said, we got together, spent the day in a studio.
(05:37):
So I picked out the cases without communicating to either Nico or Michael what they were going to be, and then shared the files with them as we went along. I mean, obviously they're all past cases and Michael has been a competitor for years, so he had seen them all at some point, but well, we can talk about how much time you spend practising for a financial modeling World Cup rounds other than doing your job. But I happen to know that that's not very much.
Host: Paul Barnhurst (06:01):
So Michael, can you share a little bit more? Anything you want to add to that or any of your thoughts on the setup and the general?
Guest 1: Michael Jarman (06:07):
As Dim said, it was set up to emulate one round from each of the competitions. And I think there was quite a good opportunity to test different things. So looking at the FMWC round, the short round was very much like how quickly can you do some back of the envelope maths that you might do when looking at a deal for the first time before building a full model free, just do the main numbers, add them up and see is this directionally going to move in the right way? Whereas the long model was like, build a full LBO model, you have an hour and a bit. So what did we have? On hour 20, I think. Yeah. And then the medium case, obviously somewhere in the middle. And then the Microsoft Excel World Championship case, that could be anything. We can go into a bit more detail what it was in a minute, but that's more to test less finance knowledge and more sort of how good are you at using Excel in a creative manner.
Co-host 2: Ian Schnoor (07:03):
And for the financial model and World Cup cases, how much of it was building from blank scratch versus for those who don't know the Microsoft Excel World Championships, correct me if I'm wrong, but you're given a decent amount of structure. And we found on the Mod Squad that it performs pretty well. AI tools are performed pretty well when there's a one example and there's a column and there's one sample. Whereas for the Financial Model World Cup, was there some structure provided or is it like, here's a blank spreadsheet, go?
Guest 1: Michael Jarman (07:34):
Yeah, so for the FNWC, 95% of the time, and I think certainly all of the questions we had here, you're given a sheet with all of the inputs on in a useful form so you don't have to spend time copying and pasting them out of the PDF. But other than that, you're just told to go. The first question, the easy one, the back of the envelope one, you're not given anything at all. Oh no, I think this one you were giving some inputs, weren't you, Dim?
Guest 2: Dim Early (07:58):
It was just an Excel file that had whatever the text of the question said, last 12 months EBITDA was X and multiple is Y and blah, blah, blah. And so the Excel sheet just wrote those numbers in cells, just like you said, so you weren't copying and pasting.
Co-host 2: Ian Schnoor (08:14):
Okay. But just so people know, some of it was kind of filling in boxes, but other parts of the simulation were building from blank scratch.
Guest 1: Michael Jarman (08:23):
Yeah. And certainly the LBA model, you're giving the inputs and that's it. You then need to start a new worksheet and build a model from scratch. So yeah, it is a pretty. Although it's still simpler than the kind of models you see in real world because you need to be able to do it in an hour and a half, not a day or a week. So it is going to be simple and have some of the sorts of annoyances of reality taken out. It has all of the meat and the potatoes that you would see in a basic LBO deal.
Co-Host 1: Giles Male (08:54):
I've got a question. So there is a full, is it two-hour raw uncut video that Nico's released? Is that right? I need to watch the whole thing. There's an audience. Who's in the audience? Where did you find the audience for this?
Co-host 2: Ian Schnoor (09:07):
What are you talking about? There was a thousand people there, gels. It was in the - You
Co-Host 1: Giles Male (09:10):
Weren't
Co-host 2: Ian Schnoor (09:11):
There. I got a custom invite from Nico.
Co-Host 1: Giles Male (09:13):
I didn't.
Guest 2: Dim Early (09:18):
No, the audience was sort of friends of the studio. It wasn't kind of financial modeling people. I'm
Host: Paul Barnhurst (09:26):
Going to have to email them. I though I got an invite. That's fascinating.
Co-Host 1: Giles Male (09:30):
And the second question I've got is for you, Dim, because for anyone that hasn't seen this yet, I think you had, I'm guessing, one of the hardest commentary jobs. I haven't watched the latter part of the two-hour video, but I'm guessing if it's literally Michael just head down and Nico going like, "Right, build this thing." And you just sat there going, "Oh God, what do I say?"
Guest 2: Dim Early (09:50):
Well, so it was interesting. I mean, we can get into it in more detail, but I mean, it was tricky to commentate on because on the one hand, FMWC was never built to be a spectator sport. And whatever, in an hour and 20-minute build, the first five to 10 minutes is mostly reading and setting up a timeline and doing some basic grunt work that is not exciting to watch at all. And there are whole stretches that are like, okay, now I've done this bit. Now the next section, okay, I need to read a bunch of technical details about how exactly this is being modeled and what the structural assumptions are and so on. So when Michael was up on screen, it was often like, oh yes, he's in the middle of doing this thing.
Host: Paul Barnhurst (10:34):
Well, you're saying it was riveting TV.
Guest 2: Dim Early (10:36):
But on the other hand, while shortcut was running, it was typically in the last maybe 5% of its runtime that it would be like, okay, now I will output boom. My sense is that structurally, the way it ran was it would model it in Python or some other under the hood coding. I think it was Python first, satisfy itself that it had the answers.Because the way Nico described it actually outputting to the worksheet in terms of effort and tokens and everything was in many ways the hardest part. Both in the sense of just writing everything to the grid, but also, okay, so what formulas are you going to use? Because part of the prompt and the challenge was don't just calculate the answers and write them all in, but build a live model. And so yeah, a lot of the time it's like, okay, we'll switch to Michael.
(11:34):
He's reading. Not knowing
Guest 1: Michael Jarman (11:35):
Very much.
Guest 2: Dim Early (11:37):
He's going to need a minute. This paragraph is quite dense. Okay, we'll switch to shortcut.
Guest 1: Michael Jarman (11:41):
Its
Guest 2: Dim Early (11:42):
Wheels are spinning.
Guest 1: Michael Jarman (11:43):
Yeah.
Guest 2: Dim Early (11:43):
There may be something short.
Host: Paul Barnhurst (11:45):
All right. And I have a dook to tell now.
Guest 1: Michael Jarman (11:49):
Going into the day, I definitely was like, hang on, I've drawn the short straw here. I've got to spend two and a half hours against a cold steel machine that's going to thrash me and Dima just has to talk about it. But then looking at some of the crowd who maybe weren't the most enthused by watching competitive financial modeling and him trying his best to enthuse them into thinking is interesting. I was like, they're either doing this actually.
Host: Paul Barnhurst (12:12):
You're like, maybe I did get the better end of this.
Guest 2: Dim Early (12:16):
Yeah. Yeah. I mean, it was interesting trying. I started off, I guess, commentating more like you would at a normal event where the audience or people who've selected into being there and they're people who do this. And then after the first one I realised, hang on. I don't think they understood most of the terms that I just used or what the question was asking. So then it was like, okay, let's talk about what a leverage buyout means and what are the kind of things you need to think about in this context and try to frame it a little bit first.
Host: Paul Barnhurst (12:49):
There was some education to the audience to help them understand what was going on. All right, so an LBO is like this, and that makes sense. And that probably made it a little easier to talk to it because a lot of the time there's nothing to talk to.
Co-Host 1: Giles Male (13:02):
And the outcome from all of this for anyone, again, that I'm guessing most people won't know, is that Shortcut won three out of four challenges against you. Is that right?
Guest 1: Michael Jarman (13:10):
Yeah. So I won the short snappy challenge. Watching it back, I'm wondering if Niko was dawdling a little bit, but yeah. I'm not sure if you
Guest 2: Dim Early (13:20):
Realised how quick you were going to be.
Guest 1: Michael Jarman (13:22):
Maybe a bit of that as well. I think he could
Guest 2: Dim Early (13:24):
Have wrapped up sooner, but he was also doing some checks and things.
Guest 1: Michael Jarman (13:28):
Yeah. So the first one, yeah, I won narrowly. The second one was like a medium level case. I though I remembered more of the case than I did and ashued reading the instructions, which I know any of the competitive people watching this know is a key element to all of these questions. So I just got stuck in a hole. I see everybody
Host: Paul Barnhurst (13:50):
Smiling.
Guest 1: Michael Jarman (13:53):
It finished in 10 or 12 minutes, which I don't think even if I'd have been perfect, I would've beaten. Then the third case, which is the full LBA model, I finished in about an hour 10 and got everything right. I think Nico pushed the bell about 35 to 40 minutes, but I think in reality it finished in 10 and then spent the next 30 checking. Because the format was that if I got 100% in an hour and he got 90% in 10 minutes, I won because I'm correct. So Niko was like -
Guest 2: Dim Early (14:25):
It probably should be, right? Yeah,
Guest 1: Michael Jarman (14:27):
Absolutely.
Guest 2: Dim Early (14:28):
Great to go twice as fast, but not if you introduce mistakes.
Host: Paul Barnhurst (14:30):
I mean, I think there's a real lesson there in the AI can build something quick, but if you want it to be right, often the majority of the time it is in checking it.
Guest 1: Michael Jarman (14:38):
Yeah. I don't think it changed anything from the checks. And there's a lot to be said for the fact that for both of us, and I got an advantage from this as well, all of the questions, all the questions are multiple choice. And the multiple choices are set. So the answers are like, is it 100,004, 100,005, 100,006? If you get minus 12,000, you know you've done it wrong. And both of us can use that. And so the AI was pretty sure. And I think Nico tried to spin off a thousand subagents to try and check the work. And Claude said, no, that's stupid. You can have 20, which is still overkill.
Host: Paul Barnhurst (15:16):
That's a lot of subagents.
Guest 1: Michael Jarman (15:18):
Yeah, it's an interesting point actually, because before we were going on stage, Nico was saying, essentially I'm going to have all of these sub-agents spinning off and checking my work. And I raised a philosophical point of how fair is that? Because I don't get to have four other mes checking my work. But I suppose in reality, there aren't four other mes. There is only one of me. So it's a fair comparison. You can copy the thing 10 times, and all it does is cost you. It doesn't cost you 10 times as many credits because all of the sub-agents work exclusively in Python. So I think they're actually quite cheap.
Host: Paul Barnhurst (15:52):
Yeah. And you can spin off lower models in sub-agents for checking, depending on the work, even lowering the cost.
Guest 1: Michael Jarman (15:58):
Absolutely. Yeah.
Co-host 2: Ian Schnoor (16:02):
Surely it does get expensive over time for him to be spending a couple hours pushing through to develop. And there's a decent number of tokens he's using up presumably in the two hours.
Co-Host 1: Giles Male (16:13):
Didn't you get a cost estimate from him or did you get a rough number?
Guest 2: Dim Early (16:17):
The way it tracked its usage was somewhat granular in the sense you could see what did this prompt cost? So there wasn't a kind of, here's one thing across all of it that we could just look at. But ballpark, I think for the big model, it was something like $20 to do the model. And then it was something less than that for many so-called adversarial review agents. Again, I think this is linking back to the idea that the expensive part was converting it into Excel, both in the sense of converting it into Excel formulas and putting it into the grid. And so the review agents would just write code, get the answer, say, "Okay, yes, I matched this answer. I am satisfied." And so they were much cheaper to run.
Co-Host 1: Giles Male (17:04):
That's really interesting. I didn't know that, that the review process is not reviewing an Excel model, it's reviewing the output of the code the builder has built.
Guest 2: Dim Early (17:14):
I mean, I think, again, I didn't get all the way into the details, but I think there was a mix of things. So there were some that were like, "Start from a blank slate, build this out. Do you match the answer?" He had prompted in various different ways. And again, he was thoughtful about using different models and at different levels of thought depending on the role. So I think for the big one, he had Fable in an orchestrator role, and then various sub-agents doing the work that were a mix of Opus and GPT, and then Fable doing some reviews and that kind of thing. So there's a lot you can think about in terms of, I mean, you can run the entire thing on Fable extra high thought, but it's going to be pretty expensive.
Host: Paul Barnhurst (18:07):
And then you wonder why you're out of tokens in 20 minutes.
Guest 2: Dim Early (18:09):
Yeah, exactly. Which I've done. But yeah, I was surprised at how relatively cheap the review agents were. Sorry, I was saying there were some that were just from scratch, but there were also some that were directed to probe a particular schedule or to review formulas in some way.
Host: Paul Barnhurst (18:32):
So he would prompt each of the agency spun up and tell them to do something. Is that the way to think of it? He says, "Spin up an agent and check this schedule or validate the code or whatever."
Guest 2: Dim Early (18:42):
I think he was mostly not giving that level of direction. I think he was giving a level of direction that was like, "Many review agents, have some check the answers, have some check the gnarliest looking parts." Again, that was the what are the bits that are most likely to trip us up or get something wrong was one of the questions put to Fable, for example.
Co-Host 1: Giles Male (19:03):
Was he asking, or sorry, was he setting that up from the start of the clock or was there pre-setup with the agent team? Do you know?
Guest 2: Dim Early (19:13):
So a bit of both. So his initial prompt, I think for all of them, or at least for most of them, was solve this. There were a few instructions that he had obviously acquired from hard experience because he had tested on a few FMWC cases previously. And he had mentioned that in one case there was something about FX rate data and there was assumptions in the file, but it had gone off and done internet research on this and that FX rate and blah, blah, blah. Gone looking for a bunch of stuff. And so he had explicit instructions about don't use anything that's not in the file. You need to compute the answers exactly. Use code to verify, spin off adversarial sub-agents to review. So he did that upfront, but then also with the LBO one, I think it finished in 10 minutes or so on the first pass.
(20:11):
And to me, the bit between 10 minutes when it finished and whatever, 35 minutes when he hit the bell, that was the most interesting time for me because that got at what is this going to look like? Okay, you can run an AI model, boom, here is the answer. What do you do between then and committing, making the decision based on that or whatever that will give you confidence that that is truly correct? And he was quite open about it. It was like, okay, well, I can have. First time around, I think he had one GPT review and one Opus review. And then he was obviously feeling a little nervous and I guess also knew that Michael was going to take a while as any human would obviously. And then he was like, okay, 20 adversarial review agents. And that was also when I think he asked for a thousand and Claude was like, no, you're not getting a thousand things.
(21:12):
But that bit of like, okay, I mean, obviously it's what it would take for him going up against Michael. He was not practically going to be able to review it himself in a way that could compete. And so in some ways that's quite different, you would hope, from a real world user who is hopefully a good enough modeler to do it themselves, but using it to save some time. And so that process of how do you think about the review? To me, that was the most interesting part of the day was how do you get from something to I feel confident in this? We talk quite a bit about during the session and afterwards. I think it's the big question. I mean, personally, I'm not thrilled about it because I am not a model auditor by background. And frankly, building model brings me much more joy than reviewing models.
(22:04):
I guess having seen this demonstration, it feels like the writing is more clearly on the wall than it was, at least for me before that, that the building is not going to be my primary value add going forward.
Host: Paul Barnhurst (22:18):
I have a question here on the build. So it got all the answers right, but as far as layout, structure, logic, how do you feel that of comparing the human? I'd love to get both of your thoughts on that. Then I'll let Giles and Ian ask any questions they have. But as you compare the two, because I've definitely seen times where it's quick, it may even be right, but sometimes the structure, I just cringe at. Sometimes it's pretty good. Sometimes it's pretty bad. So I'm curious, and you're laughing a little bit, so I'm sure there were a few things there. I'd love to get a little bit into
Guest 2: Dim Early (22:48):
That part. I mean, it was interesting. Obviously we had a pre-agreed framework for deciding who wins is like, okay, who got more right? And if they both got it all right, it was on time. And as it went, shortcut got full marks on everything. Michael got full marks on three out of the four. And the only reason he didn't get full marks on the other one was because he got to a certain point. Like he said, I think possibly spooked by how quickly shortcut went on the first one. But he was obviously like must go at full speed on the second one, bypass reading some of the instructions, built a yearly model for a question that required a quarterly model. Got to the point where he realised, okay, I can't go on without rebuilding this to expand to quarterly. And by that point, shortcut had been finished for 10 minutes.
(23:33):
He was like, I can see this round. But I have no doubt that Michael would've got all the answers on all four of them. I mean, he got all the answers on three of them right and the one he didn't, he could easily have.
Host: Paul Barnhurst (23:43):
The other one was reading comprehension, which we've all done.
Guest 1: Michael Jarman (23:45):
Yeah, yeah.
Guest 2: Dim Early (23:47):
Yes, exactly. So the interesting thing to me was if you had a sort of subjective judging element, I think the one that Michael most clearly would've lost is the one that he won. You know what I mean. It was that back of the envelope question where Michael had done it very much in a back of the envelope style like, okay, I need to divide this number by this number. Here's the first open cell. I'll just type in here, that divided by that. Okay, now I need to do something else. I'll just type that into. It was just chicken scratch. We've all done it at one time or another, hopefully to be tidied up later, but we've all done it. And Shortcut had this very nicely laid out, everything labelled, everything crisp. It built a model. And I think the reason that Michael was able to beat a computer on speed, that's kind of wild.
(24:40):
That was not where I was expecting the human to have the edge. But the reason he was able to beat it was because I think there's that kind of element of judgement that he was able to see right away, okay, I need just these three data points. I can extract them most efficiently this way. Whereas it was being quite structured about, okay, I'll build this.
Host: Paul Barnhurst (24:58):
He was being literal structured to building the whole thing to get to the bottom answer Michael saying, I'm going to take a shortcut. I need A, D, and F. I don't care about B, C, and E, so to speak.
Guest 2: Dim Early (25:08):
But having said that, I mean, you understand the value of a human doing it the most efficient way, but if you're going to have a computer do it and need to be able to review its work, you kind of wanted to lay it out that way. So yeah, my sense was it was more verbose than I would've liked in some places and bigger backup workings and that kind of thing. But then again, I'm probably stylistically less verbose than most people would like. So maybe that means it got it exactly right. But generally speaking, it produced things that looked like good models that had clearly labelled sections, formatting. They looked clean and professional. My ingoing assumption, I figured Michael was cooked on the problem solving one because I feel like that's quite, and I'm slightly sad to say this because that's probably my greater area of strength, but that's quite sort of similar to the kind of coding problems that these models have been very, very heavily trained on.
(26:19):
And it did some things that were not.
(26:27):
As a purist, I didn't like in the sense that, for instance, there was a question about filtering out only the values that were prime numbers. And as it happened, the inputs could only be numbers up to 25, and it just did a match of a list of all the prime numbers up to 25. And to me, I'm like, come on, you can write a function that will test primality that could theoretically take in. But it wasn't interested in that. It was like, let's get to the answer. But it was correct, it was efficient, whatever. Where I thought it would fall down, the one that I though Michael was most likely to win was the LBO one because it's that sort of scale where the limits of the context windows and that kind of thing can start to bite. And my sense is that that is where these models have struggled more before is on the bigger things, on the hairier things.
(27:19):
And so that was where I was expecting humanity to make its last stand rather than on. I underestimated Michael's zippiness on the short but yeah, I was not expecting it to get full marks on that question. So that was the part that most surprised me to the upside.
Guest 1: Michael Jarman (27:39):
Yeah, I think from my point of view, Dim's completely right. The models it made looked nicer and it had a spare hour, so it could spend some time making it look prettier than I did. The LBO model that I did, it had labels and stuff because I needed to track what I was doing and it was sufficiently complex that you can't do a chicken scratch model anymore. You need to label your rows so you can go and refer to them and not mess up. But it still wouldn't. If I handed it to my boss, he'd be like, "Why is this a mess? Formatted nicely." Whereas I think Shortcut did a bit of a better job on that. The thing that I didn't look too much into the model on the day, it would be good too, but I've generally found sort of, I've tried using Shortcut the last couple of days, is more on the formulaic stuff.
(28:23):
There's some things that does where I'm like, "Why have you done that?" It just makes some really strange decisions. And I think I'm not going to sit here and say, "Oh, I made a model that looked pretty, but was completely wrong because obviously you got all the answers right." But I think definitely my model, although it looked uglier, I would say had more natural and human-like formula construction than what you'd see BAI do. Certainly in the second case, which I unfortunately didn't finish, looking at some of its final formulas for the last couple of questions, it was just a bit of a mess. It kind of just brute forced one massive formula, whereas a human would err against doing that. So yeah, there's definitely some decisions that it takes and you go, "Sure, this works, but it looks weird." But yeah, it certainly. Yeah, it looked nice.
Host: Paul Barnhurst (29:17):
And it was correct. In trials, we saw a lot of that in our testing, especially early on where I saw stuff where it was linking to a number on another sheet and it used the LEP formula to link. And all it needed to do was it equals sheet A1, and it literally had a long LEP formula, and it was right. But what human would ever do that? The answer is pretty close to zero, if not zero.
Co-Host 1: Giles Male (29:41):
I've been reflecting on the whole MultiSquad series and how we've tested it, and if I can share an opinion and then just get everybody's thoughts on it. I've heard a lot about this kind of agentic or sub-agent team thing. There's somebody in the MEWC chat that's been talking about in project finance They've built a 40 sub-agent team that does all this incredible model build and review work. And I feel very strongly with my limited understanding of this that there isn't a halfway house now. What we were doing last year with ModSquad was you take a case like an FMI case and you go into Copilot and you say, "Build me a three-statement model." And for 95% of people using Copilot in a finance team, I don't think you're doing that. You're using it with Outlook and Teams and all of that great stuff. But for the, maybe it's like 0.1%, I don't know, a very small percentage of people that are really serious about modeling, why would you not have a hundred team sub-agent harness?
(30:44):
That to me is where it has to go surely, because to your point, Paul, all the things we were seeing last year with it uses a crazy LET function, you just have at least one reviewer in there that's specifically targeting don't overuse LETs or Lambdas or crazy things. I don't know. All the bit in between the Copilot studio, great. But for the financial modeling community, it feels like that's the only answer surely, that you've just got a team of agents doing this stuff.
Guest 1: Michael Jarman (31:16):
I think if you do that, I think you do need to have a large chain and it does need to be trained to really look for some really silly things. And I'll give an example. So I, in the last few days, are you Shortcut? There's a new model type that we need in my company and it's been on my to-do list. It's like, you know what? I'm just going to point Shortcut at it and I'm going to see what it does. A, to save me some work and B, just out of curiosity because I've not used it in anger before. And it did many things, but one of the things it did was it essentially had a revenue sheet, a cost sheet, a tax sheet, finance sheet. And I was like, oh, I didn't tell it not to do that. Fair enough, that's on me.
(31:51):
I was like, can you just put all of these nine sheets into one sheet, please? And you can see it's like internal thought process like, oh, this is a massive re-orchestration. We're going to have to redo all the formulas and make sure they link correctly. And I stopped it. I was like, why don't you just cut and paste it? And it was like, the user suggested cut and pasting. I can see why they suggest this, but the move to thing doesn't work. Let me look into the documentation. Oh wait, no, I can just cut and paste it. And I'm like, just some of the intuitive things where a human would just obviously be like. And it did the same thing. I was like, oh, can you space out this input sheet? It cut and paste it all to the bottom and then started cutting it all back.
(32:30):
And once I teach you what cut and paste was, cutting it all back as opposed to just adding some rows where I wanted some gaps. It works in such a non-intuitive way. And in some of the ways it does that. I'm like, this is using many more tokens than I would ever need. I could have just done this myself in 20 seconds, cut and paste nine sheets into one shin. It's taken you 20 minutes of thinking about it. So I don't know whether that's having a master fable agent at the top being like, don't be dumb. But I think indeed, because there's some things it does in an inhuman way. And sometimes it's fine, but some of the things just seems massively more inefficient than if I were to do it myself and it's slower than if I were to do it myself.
Guest 2: Dim Early (33:13):
That's a fair point. The other thing that's interesting for me, a couple of things. And again, it all comes back to the how do you get the confidence that the answers are right? Part of it is, I mean, you guys have made the point before, and I think it's still true that we are in perhaps the most dangerous time now because the models will confidently say things that look plausible. And if you don't really kick the tyres, then you're going to be in trouble. And presumably as we go forward, it will only become more likely to be correct in what it confidently says. And as you go back, it would look obviously wrong. So you'd know to check it, but this is the dangerous spot in between. But I think about this a lot of what does it mean if it is 80% right? What does it mean if it is 95% right?
(34:05):
Even at 80%, I think most people, if you tell them, "Hey, try building this with AI, but be really careful. It makes mistakes sometimes." You have them build a big model that would've taken them an hour or two or a day or whatever, and it builds it in 10 minutes and they spend an hour or two reviewing it and they find that it is completely correct. And then the next time they have it built and it builds it in 10 minutes and they spend an hour or two reviewing it and it is completely correct. I think most people, before they've even finished that second review, have already formed a mental model that says, "This is right. Why am I wasting so much time? I could build my models in 10 minutes instead of in 10 minutes of building in two hours of checking." And it's totally wrong.That thing could have a kind of 50% error rate and you just got a little lucky or it could have an 80% error rate or an 80% success rate or whatever.
(35:03):
And at the point where you get to it is by whatever way you think about the metric, it's 95%. Very few people will have the energy to really kick the tyres on that. And I think that's quite scary because just for example, and I talked to Michael about it afterwards. I think he took the same shortcut, but there was a thing in the LBO model, there was some language in the assumptions about the amortisation of the debt fees got accelerated if they were paid off in full in the period. So any remaining fees to amortise were dropped fully at that period, but otherwise it wasn't accelerated. And as it happened with the actual assumptions given, the debt did not get fully paid off in the model period. And so you could just ignore that and you'd get the answers right. And Michael made a human judgement to say, "Well, I know that there's just one set of assumptions here and this does not have to be revised and expanded and whatever.
(36:06):
So I'm going to ignore that."
Guest 1: Michael Jarman (36:09):
Shortcut kind of you. I forgot it was there. Okay, well fine. I was looking for it when I got stuck, but yeah, I forgot about it. You're being really generous, but I felt a bit mean.
Guest 2: Dim Early (36:21):
That was not in the shortcut model. And so if you change the assumptions such that the loan did get paid or revenue grows a little faster or whatever, then it would become incorrect. And for me, that's the kind of thing. I've looked at some of the headlines of these studies of how do people actually fare of reviewing AI work? I was like, if it mostly looks right, the amount of effort you put into reviewing it drops and drops and drops and drops. The chances of somebody noticing that nuance, that edge case that doesn't actually come up in this model scenario as. I mean, there are professional reviewers who are very, very good at their jobs and would, I'm sure, but most model builders are not professional reviewers. And that's the kind of thing that feels like for a lot of people actually, if you're building it in what now feels like slow motion yourself, you get to that bit and you think about the assumptions and you think about how it works and you think, oh, actually there's an edge case here in a way that if you're reviewing and you're just kind of, okay, this looks sensible, this looks sensible, this looks sensible, you probably won't.
(37:34):
So I mean, I realise that at the point where it's 95% right or 99% right, for most decisions you should probably just explore enough to get a level of comfort of when it's mostly right and go with it. Even for high stakes scenarios, a lot of models are just built as the deal partner or the MD or the CEO or whatever has a sense that about here is the right place to be. And we want to build a model to make sure we're not crazy wrong. But the model is not making the decision. The model is a sense check. And even in that scenario, the stakes feel much less high. I think in a scenario where you are truly the answer to the model is really going to drive an important decision, you probably still want to have a human build the model from scratch.
(38:28):
And maybe you want to have 50 adversarial review agents pick out it and comment and question the assumption and whatever. But I do think there's something that you get if you build the entire thing yourself that is hard to replicate in review mode.
Guest 1: Michael Jarman (38:43):
I think there's a lot on the point about amount of effort it takes to check something. So my previous role at Operus, I was head of modeling for quite a while before that worked on model audit as well. So it was one of the professional reviewers, Dim Zed. And when we're doing a model audit, that is we're giving a thing that it's 100% right. It's not 99, it's not 99.9, it's 100%. And sometimes we would be pricing a bit like a model audit and a model built together. We'll build it and we'll have someone else in the team who's unrelated audit it. And a lot of the times it'd come out the audit would cost more than the build. And we go, "Well, we can justify this to ourselves because the audit is more work." And the audit is essentially building a shadow model and then having very senior people poke at that shadow model and said it's right, so we're happy it's right.
(39:32):
But to someone looking at it who's maybe less experienced, they're going to go, "Well, hang on, why am I paying more for the audit of this model than I am for the build of it? This makes no sense at all." So I can very much agree and empathise with Dim's view as the amount of effort it takes to check something for people. It is more than you think to be 100% happy it's right and people are going to cut corners. And on that, I think myself, I have a lot of cognitive dissonance because on Excel side, when I see these AI Excel models and I spot errors in them, I'm sort of worried that people are using this and not checking it. But then when I move into something like Power Query, which I can't do, I can open Power Query and point at buttons and then ask Claude why it's not working.
(40:17):
I'm just copying the code in and going, "It's done this. Now what?" And I have no idea. And you'd think that my Excel brain would go, "Michael, this is a stupid idea. Think about all the things you know about and how wrong it is." But even then, I still do it anyway. So yeah, I can definitely imagine a lot of people very quickly just go with the Excel modeling, just go, "Oh, I must be right."
Host: Paul Barnhurst (40:38):
It's an interesting time. You've been pretty quiet in. What's your thoughts?
Co-host 2: Ian Schnoor (40:42):
Yeah, I mean, listen, I think this has been a cool experiment. I think it's interesting. You guys got some nice profile. I had a lot of fun with it. I struggled when I learned about it and I struggle at the end of this hour to know what the purpose was of it a little bit of doing it. I mean, it strikes me a little bit like when back in the day you would hear about promoter having a hundred-meter dash with a world champion, hundred metre sprinter against a cougar. Well, obviously the cougar's going to win. It's not about speed. Okay, so shortcut was faster than you. I mean, if I asked you to add up a hundred numbers on a piece of paper and Giles got to use a spreadsheet, well, guess who's going to win that as well? I'm not sure what the point was.
(41:27):
And you're one of the fastest in the world, so imagine somebody else. And I wonder whether we all have a responsibility to help people think about the purpose. Yes, we should use these tools. I mean, anyone responsible? No, why wouldn't you use these tools? Why wouldn't you use a spreadsheet? If you're still trying to build spreadsheets by hand, you don't have much of a career. You stopped working in finance in the 80s probably. So of course we should use these tools, but we're in this funny time. And then you talked about Dermond of what's the right. Well, if any associate, like when you said Dermot, people will get comfortable and not want to do their audit checks. I don't know about that. If you had any associate on your team or analyst build a model, you'd check it. And if they did 10 models, you would check them all because occasionally there might be a mistake.
(42:15):
And I think people will need to know the same thing. I mean, you don't know what could happen and they will need to be checked. What I find though is, I mean, you all, I know preaching to the choir here, but one of the problems I have about case studies is that they test that build element in the build element, but that means the instructions are clear. Well, you all know and we knew the hardest thing is you got a messy freight deal, right? Michael, you're a week away from getting into a spreadsheet. You're crawling through garbage and clunk and you're trying to read reports and ask questions and figure out what are we building? What's the model we need? What are we going with this thing? And then the last piece of it, and that's I think why our GLC, the leaders report, people unanimously felt that, well, the build portion, yeah, I mean you're sad about that, right Durham.
(42:59):
I don't blame you because it is the most fun, but the consensus was that the build portion is the portion that will likely get relegated to AI most quickly. But heck, I mean, many corporate finance teams tell me it's days until they get to the builds days, weeks. I mean, they're trying to figure out. So I'd almost love to see a situation where the test is not man versus machine with a beautiful perfect case study, but give it a big problem with a bunch of slop and a mess and don't even tell it what it has to build and don't even know the model. And the test might not even, because it still might be able to build it, but yet Nico, whoever's running it, to show up in front of a board and actually talk through and actually explain with clarity what. I mean, we all know that modeling is as much about the journey as the destination and how did you absorb through osmosis and pick up and learn and get smart?
(43:52):
I mean, listen, I love the fact that these tools are here and we should all be using them. And I feel like there's other tests that can be done, but I think that using them, finding a way to use them responsibly and elevating our game in some way is going to be crucial because people will get left behind if they don't. But anyway, Paul, going to summarise on that to say it is cool to hear you guys talk about this. And I hope we can all find a way to help those around us use tools and understand what our role has to evolve to, to still stay very highly relevant because I do believe we will continue to stay relevant if we do it right.
Co-Host 1: Giles Male (44:29):
Paul, would you mind if I just, because I assume we're wrapping up soon, can I just add my closing though to that after Ian?
Host: Paul Barnhurst (44:36):
Yeah, go ahead.
Co-Host 1: Giles Male (44:37):
Yeah,
Host: Paul Barnhurst (44:37):
Please.
Co-Host 1: Giles Male (44:39):
Yeah, I hear you, Ian. And for clarity, I'm not excited because I think we're trying to cut humans out the loop. I think the light bulb moment for me was when we are saying to people publicly, whatever, we've tested Copilot or Claude and we asked it to build a three-statement model and it was okay. It feels like that's now an inadequate test when actually there is this whole new world of capability where it can do 10 times better if you have got sub-agents and all this other stuff. But I completely agree. I think there's still humans that surely at the moment at least have to be in the loop. But the one thing I'm pleased about, I didn't think Shortcut or any other third party tool would be here. Six months ago, I would've made a lot of money that with Microsoft and Claude or Anthropic and OpenAI doing what they were doing.
(45:35):
I just didn't see a space for any other of the third parties to even survive. So I'm very happy that Shortcut, and I hope others find whatever their kind of niche is and they compete because that's going to keep, I hope, Microsoft and the other big players on their toes, which can only be a good thing.
Host: Paul Barnhurst (45:54):
So I'll share two things. I think I've been saying this for awhile and it proves what I was saying is the next step, what we tested is just seeing how good a model does with some description. But if you're working with it, it's really about have you built out instructions? Have you built out those skills? Can you run sub-agents? There's so much more than. At first, everybody's like, "I just need to be good at prompting."
(46:19):
Prompting is like 5% of this whole thing. You could have a pretty bad prompt. If you have good agents, you've given it in a lot of instruction, you have good contracts, good references, you can still get a pretty decent outcome where you could have a great prompt, but if you don't have context, the agents, the harness, so to speak is a term you hear now, which is where all the work is, where you need the expert to help build it out beyond just the AI, you're not going to get as good of an outcome. So I think what you said, Jiles, the next phase, and we've talked about that. Do we give it a workbook rule and test some of these? Do we give it some skills and see what it does? That's really the people that are getting the benefit are the ones that understand this steeply enough that they can embed all that stuff into AI and they spent the time because you get much better results.
(47:11):
And then as far as the third-party tools go, I think Shortcut's done a good job. There are several others. I've always felt like they could carve out niches, but if they're a generalist tool, they're just going to get eaten up by Anthropic, by Copilot, and McLaugh. They have to be unique. They've had to go out deep somewhere. They have to have something that makes them different for people to want to use it, because why would you spend an extra 50 bucks for something that's pretty much the same as you can get with your general subscription to whatever LLM you use? So those are some of my thoughts. Maybe I'll go to Michael and Dim, kind of final thoughts that you want to share and we'll wrap up here. We've already at, I think, 50 minutes, so we'll wrap up here in a few minutes.
Guest 2: Dim Early (47:50):
I mentioned earlier the differences between what this was testing on and real life. And there's obviously reasons that you optimise for a sort of definitive outcome in a challenge like this. But I basically agree completely with Ian. I've said for years, I've always told people you should do financial modeling World Cup. I think it's a great way, that sort of deliberate practise, the short feedback loop, the kind of do this problem for 40 minutes, you'll find out if your answers are right or wrong. You can have a model to compare to, you can try again. That short feedback loop is a great way to build execution skills. And it's hard to come by in the real world because typically people don't ask you to build models when they've already had a more experienced person, build a model, check the model, know what all the answers are supposed to be, and so on.
(48:41):
So I've always said it's great, but I've also always said that if you gave me a choice between a person who could take something that looks like an FMWC case, a perfectly structured page of assumptions, everything is nailed down to the point that you can say to seven decimal places, what is the NPV going to be under this and that assumption and blah, blah, blah, and turn that into a model and get the answers versus a person who could turn the real world with all its mess and all its human noise and crosstalk and disagreement between departments and like, oh, I think Joe in accounting has that number somewhere in a file, but he's on holiday for the next two weeks, so don't know. Someone who can turn all that kind of human messiness into, here are a set of assumptions that are credible, that the organisation is prepared to stand behind, that check out against plausible market benchmarks, whatever, that we can now use to build.
(49:43):
The second person is clearly vastly more valuable. And there are probably some parts of that, especially you think about market benchmarking, like research, AI is quite helpful with that. But there's a lot of that that is talking to human beings that I think will probably not ever be automatable because you can have a computer very efficiently send someone a questionnaire. I've been on the receiving end of these kinds of things of bank questionnaires because I'm the treasurer of our co-op. And it's like the more they automate it, the more I don't want to help them because they're trying to turn me into a computer. That versus what you will get if you go and sit and talk with a person about how they manage their data and what plausible assumptions are, and hey, what do you think about the feedback we've been getting over here is not comparable.
(50:34):
And that's an aspect of it that I genuinely think will probably never be automated and is hugely important to getting the right answer. So yeah, I completely agree with you that it's one of these things of you're measuring the most measurable thing, but that does not make it the most important thing by any means.
Host: Paul Barnhurst (50:55):
Yeah. It's a fun contest. Once a machine hits a certain spot, of course it's going to be quicker than human. We also, what happened with chess. If it gets enough structure, it's just processing.
Guest 2: Dim Early (51:08):
The one thing I will say is genuinely I did not expect that, I thought that it might win, but the one that I thought this is going to be too big, too long, too hard, I was truly surprised that it got there. Again, it's not the same as instilling the real world into that. But for me, I've made no secret of the fact that I have done very little AI. I've used AI and lots of other things, but I've done very little AI in Excel because I've followed the discussions in the MVP chats and your discussions in the Mod Squad and other things, and just had the sense until pretty recently that it just was not good enough to be worth it for someone who was excelling at my level. And now I feel like, okay, I probably do need to. Now it's good enough.
Guest 1: Michael Jarman (51:56):
Yeah, I think I agree. And I think back to my time at Operus. Operus has had to model offer champions like me and Hillary. And in both cases, what was the first thing Operus did? Promote them to a manager where they don't do modeling anymore. Now that looks dumb. And in a very small bubble it is. But then obviously you talk about the things that Ian and Dim have been talking about, the general context that they do other things. Because Operus is full of people who like to do modeling. And you look at investment banks and whatnot, a lot of associates, the first thing they want to drop when they get more senior is they want to drop the modeling, whereas that's not the case with a load of people who like spreadsheets. But it is still the case that we gave them something else to do.
(52:39):
And I look at why I'm here at my current job. Yeah, it's an advantage that the modeling I just do and it's done and it's out of the way. It's not that hard. It's one company. But the reason I'm here and why people like working with me, working with operas, working with Dim, working with whoever is what else people offer. And hopefully we don't automate that anytime soon because otherwise we are all out of a job. But I don't think we're anywhere near that point yet.
Host: Paul Barnhurst (53:05):
I agree when that happens. Let's all agree. Beach, mountain, where are we going to hang out since we don't have to work anymore?
Guest 1: Michael Jarman (53:11):
Well, yeah, well we could just have this lovely podcast and just chat all day.
Host: Paul Barnhurst (53:16):
And since everybody doesn't have work, they can all listen.
Guest 1: Michael Jarman (53:18):
Yeah, ideal. All
Host: Paul Barnhurst (53:20):
Right. Well, thank you everybody for joining. I really appreciate it, Michael and Dim, for sharing your experience in Giles. As always, I appreciate your comments and I hope everybody found the conversation useful. We did this just to kind of show you some of the things that are going on and how everybody's thinking about it. It's a rapidly changing world. I don't think any of us will dispute that. We're all nodding our heads. And the more you try to learn without letting yourself get overwhelmed, without thinking you have to keep up on every little thing, the better you'll be prepared for the future. Because the reality is the tools are here. They can help us, but there's always going to be human judgement . At least we all think there is. If that changes, we'll come on and let you know. So thanks for joining us everybody.
(54:01):
Appreciate you listening. And thanks Tim, Michael, In and Giles for joining.
Co-host 2: Ian Schnoor (54:06):
Thanks everyone. Take care. Thanks for having me.