IntellectAI · 2022–2024

First production LLM in regulated finance

I wrote the bid that won the ESG data contract for a sovereign wealth fund with over $2 trillion under management. At month ten, with the contract at risk, I proposed and proved the LLM approach that saved it.

Congratulations for delivering the first Production outcome using LLMs.

Deepak DastralaCEO, Purple Fabric, Intellect Design Arena. CTO of IntellectAI at the time.Email to the whole company, 15 August 2023
Subject line of the CTO's email, Thank you and Congratulations for delivering the first Production outcome using LLMs, sent Tue, Aug 15, 2023. The full email is in The proof below.

What happened

The arc

Machine learning took ten months to reach its limit. The replacement took nineteen days from that limit to running live.

  1. Sep 2022

    Bid won

    I wrote the tender for a sovereign wealth fund with over $2 trillion under management. Four competitors. The public procurement record names the client.

    The public record of the $2.5Mn award, with the client named

  2. Month ten

    The BERT models reach their limit

    Ten months of hand-labelled examples, and accuracy was still short of what the fund needed. The contract was at risk.

  3. Mid-2023

    Two weeks of LLM research

    A week alone, then a week with Satish Kandru: prompts, retrieval and checks tried and measured against the same accuracy targets as the machine learning approach.

  4. Aug 2023

    Five days to production

    A team of eight shipped the replacement.

  5. 15 Aug 2023

    CTO writes to the company

    Intellect's first production outcome using LLMs.

  6.  

    The fund's first commendation

    Twelve of thirteen questions taken into production.

  7. 2024 on

    Starting point for how Purple Fabric reads documents

    The rescue framework shaped how the platform finds answers in documents.

The problem

Most of the portfolio was skimmed or never reviewed at all.

The client is a sovereign wealth fund with over $2 trillion under management. It owns shares in 9,000 companies and has to judge how each one behaves on the environment, on its people and on how it is run. That is what ESG stands for. The evidence sits in thousands of pages of company reports. The brief was to read those reports for every company in the portfolio and answer the fund's own set of questions about each one, with each answer pointing to the page it came from. The public procurement record on the timeline above names the fund.

Ratings agencies could not fill the gap. Their standard products did not answer the fund's own questions, on topics like water and biodiversity, and where they had data it did not cover all 9,000 companies. The numbers show how much of the portfolio went unread.

9,000companies in the fund's portfolio
60–80reviewed in depth a year by its analysts
<10%of what the fund needed came from ratings agencies
12–18months old by the time ratings data arrived

The turn

The old models had to be taught each question from hundreds of examples. GPT-4 and Claude 2 could read the report and answer it, and with the right harness and flow around them, do that for all 9,000 companies.

For ten months the pipeline ran on BERT models, an earlier kind of language model that has to be taught each question from hundreds of examples labelled by hand. That works for questions with a clear pattern. It failed on greenwashing, a claim that sounds green but is not, because a greenwashed passage uses the same words as a real one. Those were the questions the fund cared about most. After ten months, accuracy was still below what the fund required. The table shows the two approaches side by side.

Then, in 2023, GPT-4 and Claude 2 arrived. They could read a report they had never seen and answer a question about it directly. On their own they were not enough. They needed a harness: the software around the model that gives it the right passages, asks the question the right way, checks the answer, and stops it guessing. With that harness, the same question could be asked of every report from all 9,000 companies. I proposed we build it.

To learn one question
Hundreds of passages pulled from company reports, each labelled by hand: answers the question, looks like it does but does not, or off the point
The question in plain English, plus the relevant passages from the report. The hand-labelled answers are now used only to check the model.
Judgement calls
We could tell a real answer from greenwashing. The model could not, because a greenwashed passage uses the same words as a real one, and we never cracked that.
The model reads the passage the way an analyst would, when the harness feeds it the relevant passages and checks the answer.
Time
Eighteen months of building
Two weeks to prove, five days to ship
Accuracy
Below the accuracy the fund required. No questions in production.
Met the accuracy the fund required. Questions in production across the whole portfolio.
People to run it
20 data scientists and 18 ESG analysts
18 ESG analysts and 2 data scientists. The rest of the data science team moved on to build Purple Fabric.
Built on the ML team's foundations: their question set, their hand-checked answers, and eighteen months of learning what the problem really was.

The rescue

One week alone, one week with a colleague, then eight of us shipped it in five days.

In a regulated delivery, with a client this careful, the idea met scepticism at first, and that was fair. So I asked for one week, not a project. The first week showed promise, the second held up under a second pair of eyes, and the CTO gave us the team to ship it.

1 weekof research alone: which model, which pages, how to ask, how to check
1 weekmore with Satish Kandru, senior data analyst, once the first showed promise
8people on the team the CTO then gave us
5 daysto ship the replacement into live use

The proof

The CTO wrote to the whole company the week we shipped. The client wrote nine months later.

The CTO's email of 15 August 2023, subject: Thank you and Congratulations for delivering the first Production outcome using LLMs. It thanks Aditya and Satish for challenging the status quo with the data science teams and proving that LLMs can deliver outcomes not possible with the current models.
The CTO's email to the organisation, 15 August 2023.
The full email thread, redacted (PDF)

Almost all of the questions meet or exceed our expectations. We are already working on putting 12 of 13 questions into production.

The fund's lead ESG analystMay 2024. The first formal commendation in eighteen months of service, and the first time the fund put IntellectAI-delivered data into production.

In production

All 9,000 companies in depth, and every answer traced back to the report it came from.

  • A reference panel beside page 12 of Unilever's climate transition plan, with the passage the answer came from highlighted

    Every answer links to its source page

    Click a data point and the company's report opens at the paragraph the answer came from.

  • The ESG Edge dashboard: a 70.3% overall score broken into environmental, social and governance dimensions

    Scores for the whole portfolio, down to one company

    Environmental, social and governance scores per company, down to the individual metric.

  • 10M

    documents, 60 billion searchable passages

    Every question runs against all of them, and each answer names the report it came from.

Matched or exceeded all four external benchmarks

  • Matched the benchmark
  • Missed by us
  • Missed by the benchmark, found by us
  • SBTiScience Based Targets initiativen = 532Matched the benchmark: 18%, Missed by us: 0%, Missed by the benchmark, found by us: 82%
  • Net Zero TrackerEnergy and Climate Intelligence Unitn = 124Matched the benchmark: 56%, Missed by us: 2%, Missed by the benchmark, found by us: 42%
  • Climate Action 100+Investor initiative on big emittersn = 350Matched the benchmark: 65%, Missed by us: 1%, Missed by the benchmark, found by us: 34%
  • CHRBCorporate Human Rights Benchmarkn = 241Matched the benchmark: 79%, Missed by us: 8%, Missed by the benchmark, found by us: 13%
Four public databases that track companies' climate and human-rights commitments. Most discrepancies traced back to data the benchmarks had missed. Redrawn from the team's chart, rounded to whole percentages.Original chart

What it became

A contract rescue became the starting point for how Purple Fabric finds answers in documents.

Purple Fabric is IntellectAI's platform for building AI agents inside banks and insurers. Before an agent can answer anything, the platform has to find the right passage in the right document. The question-answering framework from the rescue was the initial seed for that layer. It didn't turn into that layer directly, but it set the direction.

In January 2024, four months before Purple Fabric launched, I ran Intellect's first Prompt-a-thon: three days in which 18 ESG analysts learned to build with LLMs from scratch.

The ESG team gathered at the close of Intellect's first Prompt-a-thon: 18 ESG analysts across three days of LLM training
The team at the close of the Prompt-a-thon, January 2024.