gm — some of you know Participation Architecture ( Participation Architecture - Final Grant Report ) (Arbitrum grant, S3). I’m extending it into PhD research on delegate fatigue: what makes voting cognitively heavy, and whether we can actually measure it.
I’m piloting a Delegate Fatigue Index — a tool that scores governance workload per vote — and I need 3-5 active delegates for a 45-min call:
- you go through a short survey and try the tool
- you think out loud — what’s clear, what’s confusing, what’s off
- I listen and take notes (recorded with your consent)
No prep. If you’ve voted in the last ~2 weeks, you’re a perfect fit.
Comment here or DM me and I’ll grab a slot in your timezone: https://cal.com/wyszomirski/30
1 Like
Interesting to see delegate fatigue treated as a measurable governance workload per vote. Could you share more on what theoretical framework and data you used to design and calibrate the Delegate Fatigue Index, so active delegates can better understand how accurately it reflects our lived experience? @wyszomirski
Good question, and it points straight at what the study is for.
The DFI does not claim to be authoritative on its own. It is a testable proxy. Five components, each grounded in a kernel theory: volume and concurrency treat attention like a rivalrous commons (Ostrom); burstiness and reading time come from the Fogg Behavior Model, where spikes break the regularity habits need; novelty comes from Cognitive Load Theory, where new decision types cost more than routine ones.
Reference values are calibrated on real Arbitrum data. Median proposal runs ~1,400 words. Median delays. Those anchors define"normal" load.
Calibration is not validation, though. Whether the index matches your lived experience is empirical. I do not get to settle it by design. So I test it directly: each delegate rates their own recent vote on the NASA-TLX, a validated workload scale, and I check whether DFI tracks that rating, per delegate, across subscales.
That check is the 45-minute session itself. You vote, you rate the load, I compare. If the index disagrees with active delegates, the index is wrong, notyou.
If you vote on Arbitrum, you would be a strong fit - happy to send details. If your interest sits mainly on the research side, I would also value comparing notes on multi-chain workload.
Creating a new measure for a narrow use case is generally bad OB or b school research. Why would you create a more narrow measure than existing measures that predict much more, we have measuring for cognitive fatigue, we have essentialist measure for need for congition etc. you are just taking existing theorteical findings and slicing it the constructs more narrow so it predicts less and discovers nothing new… this would not pass in core psych but maybe it does it b school….
Also, don’t dump AI slop to people if you want participation..
Fair hit on the writing. Keeping this short.
DFI is not a psychological scale and does not compete with NASA-TLX or Need for Cognition. Both of those require asking the person. DFI reads what a delegate already faces - proposal volume, overlap between proposals, burstiness, reading time, novelty - straight from Snapshot and Tally. No survey involved.
That buys coverage. It scales to every delegate in a DAO, including the ones who never answer a questionnaire. It also creates the obvious problem: an unvalidated index is worth nothing.
So the study validates it against a validated instrument. Each delegate rates one recent vote on NASA-TLX. I compare that rating against the DFI computed for that same vote, per subscale, per person. If DFI does not track the ratings, the index fails and the dissertation reports that.
NFC is a trait, TLX is a task state, DFI is an observed workload trace. They measure three different things.
If a measure already estimates per-vote governance load without asking the delegate, name it. I will test against it.