The Shape File Bridge
The shape file bridge is a working pattern for building analytical tools on confidential operational data when two constraints hold at once: the data may not leave the organization that owns it, and the people or systems best placed to build the tool work outside it. The pattern separates the structure of a problem from the identity of the parties in it. Structure crosses the boundary in a small, machine-readable shape file; identity never does. A tool is built outside against the shape file, published as a single self-contained HTML file, and carried back in, where the analyst who raised the problem loads the real data and the real names on a machine that never shared them. The pattern was developed for workforce planning teams in large service businesses, where each client or business unit supplies a differently formatted forecast or performance file and every urgent question is answered with a hand-built slide, but it applies to any analytical function that works inside a data boundary.
The problem it solves
Planning and analytics teams inside large organizations face a specific bind. The questions they are asked are varied and urgent: a client's service level is slipping, a migration is landing early, a forecast file in an unfamiliar layout has to be reconciled by Monday. The data needed to answer them is confidential by contract or regulation, so it cannot be pasted into an external service or sent to an outside builder.[1] Meanwhile the capacity to build durable tooling, whether that is a senior analyst with time to think or access to a capable model and a development environment, is usually outside the team, outside the network, or rationed.
The default resolution is the spreadsheet and the slide. Each question gets a workbook built by hand and a deck assembled from screenshots of it. The research on end-user computing has shown for decades that this is where errors live: field audits consistently find material errors in a large share of operational spreadsheets, and the errors are rarely caught because nobody tests a model that was built in an afternoon.[2] The deeper cost is that the work does not accumulate. The next question starts from a blank sheet, and the team's knowledge of how to answer it lives in whoever built the last one. The intake discipline reduces how many such questions reach the team; the shape file bridge changes what happens to the ones that do.
The principle: structure crosses, identity does not
A planning problem can almost always be described without naming anyone. A requirement curve by week, a supply plan with training and nesting waves, a daily table of offered and handled contacts by channel, a set of assumptions with ranges: none of these needs the client's name, the vendor's name, the people's names, or even the real scale of the operation to be analysed correctly. What the analysis needs is the shape of the data, its columns and their meaning, the relationships between series, the assumptions and how confident anyone is in them, and the question being asked.
The shape file carries exactly that and nothing else. It follows the logic of pseudonymisation as data-protection law defines it for personal data, applied here to commercial identity as well: the identifying attributes have been removed or replaced such that the data cannot be attributed to a specific party without additional information that is kept separately.[1] The additional information, in this pattern, is a code-name key that never leaves the organization, or simply the analyst's own knowledge of which client the file describes. The identity field in the file is either a code name or blank, and it is filled back in on the return leg, locally, by the person who removed it. The distinction between pseudonymisation and full anonymisation matters here: the file is not being released to the public, it is being carried across one boundary by one trusted person, and the risk model is sized to that.[3][4]
Structure alone can still identify. A migration of an unusual size on an unusual date, a channel mix nobody else has, or a very small population can point to one client as surely as a name.[5] The pattern answers this with scale stripping: the shape file can be exported shape-only, with every absolute quantity divided by a reference size so that only ratios and curves remain, and multiplied back on the return. Dates can be shifted to a week index. The builder outside sees that January is short by a quarter of the requirement; only the analyst inside knows that the quarter is forty people.
The loop

The bridge runs as a loop with six stages. Three happen inside the boundary, two outside, and the sixth is the crossing itself; the pack makes a second crossing on its way back.
| # | Stage | Where | What happens | What crosses |
|---|---|---|---|---|
| 1 | Problem statement | inside | The analyst opens the pack's intake tool in a browser and answers a structured interview: what decision the analysis changes, who owns it, by when, what the data looks like, what is already known and what is assumed. A model in a desktop assistant can conduct the interview, but it only ever sees the structure. | nothing yet |
| 2 | Shape export | inside | The analyst maps the columns of their own file, however it is laid out, onto a published data contract. The tool validates the mapping, grades each series, strips identity and optionally scale, and writes one JSON shape file. | — |
| 3 | Transfer out | the boundary | The shape file leaves by ordinary means: an email to the builder, or a self-expiring link. It is small, text, and carries no names. | the shape file |
| 4 | Build | outside | The builder, human or model, works against the shape file and the contract. The problem is understood from the structure, the method is chosen, the tool is written and tested on synthetic data in the same contract. This is where the heavy reasoning and the expensive compute happen, on infrastructure the organization does not have to provision or secure. | — |
| 5 | Publish as an HTML pack | outside | The tool is published as a single HTML file, with no server, no network calls and no external resources, together with the contract it reads, the instructions for the desktop assistant that will run the intake next time, and a page describing the problem archetype it solves. None of this carries client information; the pack is reusable by anyone with the same shaped problem. | the HTML pack |
| 6 | Return and re-identify | inside | The analyst receives the HTML file by email, opens it locally, loads the real data file and the real identity. The tool regenerates the view every time the data changes. The story is told from a tool, not a slide, and the next analyst with the same shaped problem starts from the pack rather than a blank sheet. | nothing |
Two properties make the loop safe. First, every artefact that crosses is inspectable: a shape file is readable JSON and an HTML pack is readable source, so either can be checked by eye or by a scanner before it moves. Second, the return leg needs no trust in the outside at all. The HTML file is built to make no network requests, and because its source is readable, that can be verified by inspection or a scanner before it is opened; a tool that has not been checked should not be trusted with real data.
The shape file
A shape file is one JSON document with a fixed envelope and a payload defined by one or more data contracts. The envelope is the part every tool shares.
| Section | Carries | Why it is there |
|---|---|---|
| Schema id and version | which contract family and which version wrote the file | so a tool knows what it is reading and whether it can |
| Scale | scaled (real quantities) or shape-only (quantities divided by a reference size) | so absolute scale can be withheld without losing the curves |
| Provenance | the tool and version that wrote the file, whether a person or a model generated it, a hash of the payload, and an explicit assertion that the file carries code names only | so two files can be compared without opening them and a file with real names is caught at the door |
| Basis | reference size, horizon, channel mix, whether supply is already net of shrinkage, the week-zero date | the handful of facts every calculation depends on |
| Assumptions | each as a value, a low–likely–high range, a status (default, estimated, confirmed), an owner and a grade | so the number can be explained and challenged, not just displayed |
| Grades | one per payload column: measured, computed, estimated or asserted | so every headline number can show how much to trust it, and computed series inherit the weakest grade among their inputs |
| Payload | the data itself, in named contracts with enforced columns | the shape |
| Questions | open questions with an owner and a status | the parts of the problem statement the data could not answer |
Three design rules keep the format usable across many tools and many years. The reader is tolerant: unknown keys are ignored, missing keys take defaults, and a file from an older or newer tool still loads, with a list of what was dropped. This is the robustness principle long applied to network protocols, and the same reasoning applies to a file that will be written by one tool and read by another years later.[6][7] Versions follow the semantic versioning convention, under which an additive change increments the minor number and an incompatible change increments the major.[8][9] Applied to the envelope, this means keys may be added freely, while renaming a key or changing its meaning is a major release; by design, a tool loads any file within its own major version and reports a different major rather than refusing it. And the validator enforces the contract rules at paste time rather than at presentation time: a service-level numerator larger than the handled count, a dataset mixing two definitions of handle time, or shrinkage applied twice is flagged before anyone builds a view on it.
The grading convention deserves a note. Marking each series measured, computed, estimated or asserted, and propagating the weakest grade through every calculation, is a lightweight form of the calibrated-estimate discipline: it does not make a weak number stronger, but it stops a weak number from being presented as a strong one.[10]
The data contract
A contract is a named table with enforced columns, published alongside the tool. The analyst's own file is never expected to match it; the intake tool maps columns interactively and remembers the mapping, so an awkward client export becomes a one-time chore rather than a weekly one. A small family of contracts covers most of the recurring views in a planning function: weekly supply and requirement by site or line of business, planned waves with training and nesting, daily performance by channel. New contracts are added when a view needs columns no existing contract carries, and they are additive, so a tool that reads one contract is unaffected by the existence of another.
The contract is also the place where definitional ambiguity is forced into the open. Whether handle time is elapsed or worked, whether supply is gross or net of shrinkage, and from which week a go-live is counted are each a declared field in the contract rather than a convention in someone's head. A file cannot be exported without declaring them, which means someone inside has to answer the question, which is often among the more valuable things the first export produces.
The HTML pack
An HTML pack is the published, reusable form of a tool built on the bridge. It is distinct from a project pack, which carries instructions and reference material for an assistant, and from a deck rebuild kit, which carries a presentation. An HTML pack carries four things:
- The tool: one HTML file, self-contained, that opens from a local disk in any modern browser, makes no network request, embeds its own styles and scripts, and stays small enough to travel as an email attachment. It reads and writes the shape file, accepts pasted or uploaded data against its contract, shows every headline number with its grade, and exports an update note that can be dropped into slide notes or a status message.
- The contract: the columns the tool reads, their meanings, and the validator rules, so an analyst can prepare data before opening the tool and a builder can extend it without guessing.
- The assistant kit: an instruction block and context files for a desktop assistant project, so that the next problem statement is collected the same way and the shape file is generated rather than typed.
- The page: the problem archetype the tool addresses, when to use it and when not to, and the download address of the tool.
The single-file constraint is a security property, not a convenience. A file verified to make no network calls has no channel to leak through; a file with no server has no server to patch; and a file that opens from an email attachment needs no procurement, no installation and no exception from a locked-down laptop policy. The costs are real too: the tool cannot be updated in place, so versions must be visible in the file and the pack page, and a tool that does arithmetic the organization later disputes is as auditable as its source, which is why the source ships in the file.
Roles
The bridge works because each side does what it is placed to do. The analyst inside owns the question, the data, the identity and the decision. They are the only party who ever sees real names beside real numbers, and they own the number that leaves the tool. The builder outside owns the method and the tool. They see structure, grades and assumptions, and they are expected to challenge them. The pack is the contract between the two: it says what the tool needs, what it promises, and what it will refuse to compute. A function that runs this way can route its heavy builds to whoever is best placed to do them, inside or outside, without renegotiating the data boundary each time.
Failure modes
- Structure that identifies. A shape file from a small or distinctive operation can be recognisable even without names. Shape-only export, date shifting and a scanner that checks every outbound file against a code-name list are the controls; the residual risk is accepted knowingly, by the analyst, per file.
- The key in the file. If the code-name key or the real identity is ever written into the shape file, the pattern fails silently. The provenance assertion and the scanner exist to catch this, and the key is kept outside every repository and every shared folder.
- Drift between tool and contract. A tool built against one major version of a contract loads a file from another only as far as the keys still match, and it must say what it dropped. Semantic versioning and the tolerant reader reduce the blast radius; the pack page records which contract version each tool reads, and a view built on a partially loaded file is marked as such.
- The tool as shadow system. A single-file tool that becomes the system of record for a decision has escaped governance. The pattern's answer is that the tool regenerates a view from a file that has provenance, grades and an owner; it is not a database, and the pack says so.
- Transport friction. Mail gateways strip or block HTML attachments. The sanctioned route is an approved file-transfer channel or an allow-list entry agreed with the security function, and the tool should stay small enough to attach. A self-expiring link is the alternative when policy allows.
- Hand-built again. If the return leg ends with the analyst screenshotting the tool into a slide, half the benefit is lost. The update-note export and a slide template that accepts the tool's panels directly are what close that gap.
Relationship to agentic planning
The bridge is the manual precursor of an automated planning chain. Each stage of the loop corresponds to a step that an agentic workforce-management protocol eventually runs without hands: intake becomes a structured capture, contract validation becomes a gate, the build becomes a deterministic engine chosen by a model, and publication becomes a regenerated view. Running the loop by hand first does three things that automation later depends on. It produces the contracts, which are the interfaces the automated steps will share. It produces a catalogue of problem archetypes, each with a tool that already works. And it produces evidence, in the form of how often the validator caught a real defect and how often a regenerated view replaced a hand-built one, that a function can use to decide which steps to automate and in which order. A natural first gate for such a programme, that a view regenerates from a file with no hands, is exactly what one completed loop demonstrates. See Agentic AI Workforce Planning and The Agentic Handover Gate for the automated end of that path.
Maturity Model Position
- Level 2: every question answered with a fresh workbook and a deck; data moves by screenshot; nothing accumulates.
- Level 3: recurring views regenerate from shape files through single-file tools; contracts exist and are enforced at paste time; the problem statement is captured before the build starts. This is the level the bridge establishes.
- Level 4: shape files are generated rather than typed, the desktop assistant conducts the intake and drafts the assumptions register, and the catalogue of packs covers most recurring views; build requests route on evidence of fit.
- Level 5: the loop runs as an automated chain inside the organization's own environment, with the bridge retained only for problems the chain has not seen before.
See also
- Work Intake for Planning and Analytics Teams, the gate that decides which questions reach the bridge
- Agentic AI Workforce Planning and The Agentic Handover Gate, the automated end of the same path
- WFM Data Governance and Quality and GDPR and Workforce Data, the governance the pattern operates under
- Deterministic vs Probabilistic Models, on choosing the engine inside the tool
- Fixed-Capacity Service Models and Pooling Architecture in Service Workforces, recurring problem archetypes that fit the pattern
- Wiki:HTML Packs, the register of published tools; Wiki:Library for the other kinds of deployable material
References
- ↑ 1.0 1.1 Regulation (EU) 2016/679 (General Data Protection Regulation), Article 4(5) and Article 5(1)(c). Official Journal of the European Union L 119, 4 May 2016. eur-lex.europa.eu
- ↑ Panko, R. R. (1998, revised 2008). "What We Know About Spreadsheet Errors". Journal of End User Computing 10 (2), 15–21. doi:10.4018/joeuc.1998040102
- ↑ Garfinkel, S. L., Near, J. P., Dajani, A. N., Singer, P., and Guttman, B. (2023). De-Identifying Government Datasets: Techniques and Governance. NIST Special Publication 800-188. National Institute of Standards and Technology. doi:10.6028/NIST.SP.800-188
- ↑ ISO/IEC 20889:2018. Privacy enhancing data de-identification terminology and classification of techniques. International Organization for Standardization. iso.org
- ↑ Sweeney, L. (2002). "k-Anonymity: A Model for Protecting Privacy". International Journal of Uncertainty, Fuzziness and Knowledge-Based Systems 10 (5), 557–570. doi:10.1142/S0218488502001648
- ↑ Braden, R. (ed.) (1989). Requirements for Internet Hosts: Communication Layers. RFC 1122, §1.2.2 "Robustness Principle". Internet Engineering Task Force. rfc-editor.org
- ↑ Fowler, M. (2011). "Tolerant Reader". martinfowler.com. martinfowler.com/bliki/TolerantReader.html
- ↑ Raemaekers, S., van Deursen, A., and Visser, J. (2014). "Semantic Versioning versus Breaking Changes: A Study of the Maven Repository". Proceedings of the 14th IEEE International Working Conference on Source Code Analysis and Manipulation (SCAM 2014), 215–224. doi:10.1109/SCAM.2014.30
- ↑ Preston-Werner, T. (2013). Semantic Versioning 2.0.0. semver.org
- ↑ Hubbard, D. W. (2014). How to Measure Anything: Finding the Value of Intangibles in Business, 3rd ed. Hoboken: Wiley. Chapters 5–6 on calibrated estimates.
