I.
Published · SSRN
What the Model Remembers
Abbie R. Lee·Andrew Koh·EigenFlow Research
The paper takes a question most people treat as a tradeoff and shows it is a single
design problem. Suppose an intermediary trains a model on data pooled from competing
general partners, then runs that model inside each member's own valuation, reporting,
and underwriting. Everyone's estimates get sharper. Valuation errors compress,
dispersion in reported NAVs falls, and the information component of secondary
discounts narrows.
The catch is that a learning system leaks. It leaks through its outputs, through its
parameters, and through the queries it answers. Worse, accuracy on a long-tailed
corpus requires memorizing the tail, and in private markets the tail a model
memorizes is precisely a manager's proprietary thesis.
What the paper establishes is the set of conditions under which that leakage stays
bounded: per-contributor influence limits, a certified class of permitted queries, a
metered privacy ledger, and attested sealed execution. Under those conditions the
exposure is capped uniformly, holds up against collusion, survives any downstream use
of the model, and shrinks as the pool grows. Without them, memorization reinstates
exactly the adverse selection that empties voluntary databases of the funds worth
observing.
Same enclave, same agents, same data. The training procedure decides which outcome
you get.
Proprietary data
Memorization
Privacy
Information production
Fund administration
II.
In preparation
Sharing Without Showing
EigenFlow Research·Working draft
The companion paper turns from what a model remembers to what a counterparty can
verify. If a limited partner cannot inspect the underlying positions, on what basis
should they believe a reported mark? And if a general partner cannot reveal those
positions without giving away the thesis, what can they prove instead?
The draft develops the mechanism side of the same architecture: what an attested
computation can certify to a party who never sees the inputs, and what that
certification is worth in a market where discretionary valuation is the norm.
This paper is still in preparation and no preprint is posted yet. If you would like
the draft when it circulates, write to
gp@eigenflow.pro
and we will add you to the list.
III.
Note
From the research to the product
The reason we publish is not credentialing. The conditions the first paper derives are
the conditions our product has to satisfy, and writing them down formally is how we
check whether the thing we built actually does what we claim.
Each result maps onto something a fund manager can see in the software:
Per-contributor influence limitsPaper I, §4
No single fund's data can move the shared model enough to be recovered from it. In practice this is what lets a manager contribute to a pooled model without contributing their edge.
Certified query classPaper I, §5
The system answers a fixed, auditable set of questions. Queries designed to extract a specific holding are not in the set, so they cannot be asked.
Metered privacy ledgerPaper I, §5
Cumulative exposure is tracked and capped rather than assumed away. Every answer draws down a budget that does not refill.
Attested sealed executionPaper I, §6
Analysis runs where neither we nor another member can observe it, and the attestation is verifiable rather than promised.
What this adds up to for an operating fund is ordinary enough: capital accounts that
reconcile, LP statements that go out on time, and marks a limited partner has some
reason to trust. The architecture is the interesting part. The output is supposed to be
boring.