Client-Ready Decks for Relationship Managers Working with Client-Identifying Data: PowerPoint Generation on Swiss-Hosted GPT-5.1

7 min read

Highlights

  • A relationship manager (RM) needs client-ready decks, and they contain client-identifying data (CID). Meeting preparation, portfolio reviews, proposals and pitch books are PowerPoint decks in the bank's template, full of client names, holdings and family details.
  • CID means Swiss-only inference, and that means GPT-5.1. Our Swiss clients' GDPR guidelines require that CID is processed only in Switzerland. GPT-5.1 is the strongest model Microsoft offers with Swiss-only inference; no stronger model is available in-country.
  • A standard PowerPoint skill fails with GPT-5.1. Frontier models manage it, but take over 20 minutes and about $8 per deck, and an RM cannot use them with CID.
  • Our approach: GPT-5.1 writes the content, a compiler builds the deck. The model only decides what goes on each slide. A deterministic compiler places it into real slides of the bank's own template and checks that every text fits.
  • Result for the RM: a complete deck in the bank's template in 2–10 minutes, for $0.14–0.77. The RM chooses and confirms the template in a short selection step.

The RM's situation

An RM prepares for a meeting with an entrepreneur family tomorrow. They want a deck with the family's current allocation, performance against the agreed strategy, the proposed changes, fees and next steps. The deck must be in the bank's template, look like the bank made it, and every number must be right. The input is CID: names, holdings, family situation. So the assistant runs on GPT-5.1 with Swiss-only inference.

This is one of the most valuable tasks an agent harness can take off an RM: a deck that takes an afternoon should take minutes. It is also one of the hardest for a model of GPT-5.1's size. A standard skill asks the model to write code that builds slides shape by shape, to keep the template's look consistent over many slides, and to judge whether text fits. GPT-5.1 cannot see the result, and the errors add up.

What the RM gets

  • The bank's template, chosen and confirmed. A template-selection skill offers the bank's templates and the RM confirms one, or attaches their own. Every deck is built from the confirmed template. A new or changed template is analysed automatically the first time it is used (7–25 seconds, no model calls).
  • Real template slides. Content goes into copies of the template's own slides: cover, agenda, KPI tiles, charts, tables, timelines, process flows, team slides, funnels. Fonts, colours, logos and footers stay the bank's.
  • Charts from the client's numbers. Bar, column, line, combo and donut charts (such as an allocation donut or a performance trend) are filled with the actual data and stay editable in PowerPoint.
  • Large content split cleanly. A long holdings table, many KPIs or a large team are spread over several slides of the same type ("1/3, 2/3, 3/3") instead of being shrunk or cut off.
  • Honest about what the template lacks. Elements with no template equivalent (maps, video, QR codes, animations, speaker notes, custom fonts) are replaced with the closest template component or left out. The reply tells the RM exactly what was substituted and what to add by hand.
  • Research and languages. Market context can be researched on the web first. The deck is written in the RM's language (tested in English and German).
  • Fast and cheap. 2–10 minutes and well under $1 per deck, on a model the RM may use with CID.

How it works

  1. Template analysis. Every slide of the bank's template is measured: layouts, text slots and how much text each holds, charts, tables and repeated elements. The result is a recipe per slide that the model reads.
  2. Content. GPT-5.1 writes per slide a title, a key message and typed blocks (bullets, metrics, chart data, table, people, timeline), optionally choosing a recipe.
  3. Compile and check. The compiler fits each block to the best template slide. It fills a copy of that slide and rejects anything that would overflow, with a reason the model can act on. A final check confirms the delivered file is the compiler's output.

The model never positions shapes, picks fonts or colours, or writes slide-building code. Those are exactly the steps GPT-5.1 gets wrong.

How we test it

Test harness. A script sends each prompt through the platform's public API into a QA space running the skill: the Codex harness with GPT-5.1 at medium reasoning effort. For each chat, it:

  • downloads the generated deck;
  • renders every slide through Microsoft PowerPoint;
  • records time, tokens and cost.
Prompt setCountWhat it covers
Reference request1A real 27-slide request from the banking industry, used to compare models and skills on time and cost
Realistic cases10Typical RM and banking material: family-office first meeting (German), pension fund ESG report, M&A board paper, fundraising deck, investor update, market landscape, regulatory briefing, QBR, earnings flash note, vague one-line brief; 4 need web research first
Adversarial cases5Requests the template cannot meet as asked: waterfall, scatter and gauge charts; SWOT, BCG, five forces, org chart; far too much content per slide; video, QR codes, custom fonts, speaker notes; off-script edits (move the title, Arabic right-to-left, copy a sample slide)

What we check per deck.

  • Delivered versus requested slides.
  • Compiler report: how many slides are template slides, composed from template parts, or rejected.
  • The deck is compiler output.
  • Visual review of every slide: template fidelity; overflowing or cut text; empty placeholders; leftover graphics; readable charts; correct numbers.
  • Time and cost.

We test on a corporate template from the banking industry (not shown) and on the Unique Company Template. All published evidence uses the Unique Company Template: see Evidence: RM Decks on the Unique Company Template (GPT-5.1). The test prompts use fictitious clients; no real CID is involved.

Results

Reference request (27 slides): time and cost by skill and model.

SkillModelUsable with CIDTimeEst. costResult
Standard pptx skillOpus 5 (high)No23m 0s~$8.09Success, small corrections
Standard pptx skillGPT-5.6 Sol (medium)No24m 19s~$8.52Success, small corrections
Standard pptx skillGPT-5.1 (high)Yes12m 36s~$0.83Failed
Compiler skillGPT-5.1 (medium)Yes4m 46s~$0.40Success, template attached

For the RM, that is a deck about 5× faster and 20× cheaper than frontier models with the standard skill, on a model that can be used with CID.

Unique Company Template as the selected template: typical RM requests.

CaseTimeEst. costResult
Family-office first meeting (German)2.5 min~$0.217 slides, none rejected; services, investment philosophy, model allocation, process, team, fees and next steps
Pension fund ESG report3.7 min~$0.309 slides, none rejected; KPI tiles, carbon pathway chart, allocation donut, roadmap
M&A board paper4.7 min~$0.398 slides, none rejected; valuation and financing tables
SaaS QBR1.9 min~$0.1410 slides, none rejected; KPI tiles, trend chart, cards
Strategy frameworks (adversarial)2.7 min~$0.199 of 9 slides; SWOT, BCG and five forces as clear bullet slides, funnel and priority table from the template
Overload (adversarial)9.5 min~$0.7711 slides, none rejected; 12 KPI tiles over 2 slides, a 16-row table over 3, 10 people over 2, a 20-step journey over 2; the reply explains the splits

Banking-industry corporate template.

  • Realistic cases: 9 of 10 produced complete decks, in 2–10 minutes for $0.21–0.86 each.
  • Adversarial cases: the model stayed inside the template in all 5 and listed every substitution. For example, waterfall charts became bar charts, heat maps became tables and gauges became metric tiles.

What the RM still has to do

  • Content is checked by the file hallucination check (beta). The compiler guarantees layout, not content. A generated deck is checked against every source the agent used, and the file card shows a warning when the deck contradicts its sources, e.g. a wrong figure, date or classification. See Hallucination Check: a second compliance layer for relationship managers. The RM still reviews the deck before it goes to the client.
  • Occasionally, fix cosmetic points: a plain title slide instead of the image cover, a text-heavy slide, or German umlauts written as "ae/oe/ue".
  • Diagrams the template does not contain (SWOT grid, BCG matrix, Venn) come as structured text, not drawings.
  • The model is instructed to always use this skill for PowerPoint. In rare cases it may still build a deck with its own code; a platform-side check of the output file would rule this out.
  • A bank's own template is analysed automatically. Each new template should get one review round before it is rolled out to RMs.

Problem statement

An RM spends a large part of their client work producing documents: meeting preparation, portfolio reviews, proposals, pitch books and follow-ups. These are almost always PowerPoint decks in the bank's corporate template. That makes document generation one of the capabilities of our agent harness that matters most to RMs. It is also where the value is most visible: a deck that took an afternoon should take minutes, and it must look like the bank made it.

This material contains CID by nature: names, holdings, family situations. Our Swiss clients' GDPR guidelines require that inference on CID happens only in Switzerland. Today, the strongest model Microsoft offers with Swiss-only inference is GPT-5.1. Frontier models such as Opus 5 or GPT-5.6 Sol produce good decks with a standard skill, but an RM cannot use them on CID in Switzerland. So for the RM's most important use case, the harness has to work with a model that is noticeably weaker at long, multi-step, visually demanding tasks.

Our approach moves the hard parts out of the model. Template knowledge is extracted once into measured recipes. GPT-5.1 only decides what goes on each slide. A deterministic compiler does the layout and refuses what does not fit. This makes the model's job small enough for GPT-5.1, and the deck's quality comes from the bank's template and the compiler rather than from the model's design ability.

The same pattern should carry over to the RM's other documents: Word reports, Excel overviews and client letters. Wherever CID pins a deployment to a less capable model, moving deterministic work out of the model and into tools narrows the gap to frontier models.

Evidence

The Unique Company Template and every produced deck, with all slides: Evidence: RM Decks on the Unique Company Template (GPT-5.1).

Last updated