Independent Architecture Review — Arbitrum Stylus Developer Impact Oracle
@JulianCross — I compared the new Stylus Developer Impact Oracle RFC against the evidence boundary established in the Stylus Mobile Native Authentication discussion.
The direction is worth pursuing. Arbitrum does need a way to distinguish funded work that produced durable ecosystem value from work that only satisfied a reporting surface.
The first blocking problems visible in the public specification are bounded, and they sit before the SQL layer.
The current public RFC has moved the indexer/dashboard one step ahead of the object-definition and evidence layer. If that is corrected first, the later implementation gets a stable target. A complete implementation review will still require the actual SQL/indexer/schema once those artifacts are public.
The most important point is also a small specification regression from the framework you already accepted on August 24.
1. First boundary: the retention clock has moved upstream again
In the Mobile Native Authentication thread, the agreed ordering was:
Funded obligation
→ Reproducible build
→ Deployed artifact
→ qualifying production evidence
→ retention
You explicitly agreed that the Oracle must ingest the funded-obligation → reproducible-build → deployed-artifact lineage before the retention clock begins, and that the ingestion layer should preserve:
ACCEPTED
EXCLUDED_WITH_REASON
DISPUTED
UNRESOLVED
SUPERSEDED
The new RFC currently starts Module A from:
cohort wallet addresses
→ Stylus deployments
→ 30 / 90 / 180 days post-grant
That re-opens the exact boundary established in the August 24 discussion.
A cohort wallet is not the funded object.
A deployment is not proof that the funded source produced that artifact.
And “post-grant” is not yet a stable retention anchor.
The indexer can therefore be perfectly deterministic while measuring the wrong lineage.
Minimal repair
Insert a Qualification & Provenance Gate before Module A.
At minimum, the canonical grant object should bind:
grant_id
program_id
program_class
funded_entity_id
approved_objective_hash
approved_scope_hash
approved_milestone_set
funding_commitment_id
repository_set
methodology_version
cohort_manifest_hash
eligibility_rule_hash
cohort_frozen_at
The cohort and its eligibility rule should be frozen before outcome observation. Otherwise the denominator can drift after results are visible.
For code-bearing grants, the qualified artifact lineage should then bind:
source_commit
build_manifest_hash
artifact_hash / WASM hash
deployed_code identity
chain_id
deployment_tx
deployment_block_hash
initialization / constructor commitment
upgrade lineage, where applicable
provenance status
Identity references should be domain-qualified rather than stored as bare strings:
AddressRef = (chain_id, address)
ContractRef = (chain_id, address, code / upgrade epoch)
RepoRef = (host, owner, repository, commit)
The same hexadecimal address on two chains is not the same observation object, and an upgraded contract is not automatically the same code identity.
Only after the relevant object reaches ACCEPTED should a retention clock exist.
The clock should be bound to a qualifying event, not to an ambiguous calendar label. The specification should also define whether “30/90/180-day retention” means activity at a point, activity inside a fixed observation window, or another declared rule; different window semantics must not share the same label:
retention_t0
=
first finalized qualifying event
defined by the grant's objective class
For one project that may be first production execution.
For another it may be the release of a production package, an upstream integration, or another objective-specific event.
The important invariant is:
NO QUALIFIED OBJECT
→ NO RETENTION CLOCK
This should be a hard fail-closed condition.
2. The Oracle needs a Program/Impact Class before it needs one universal metric
The current RFC references STIP, LTIPP, Questbook Domain Allocator rounds, Stylus deployments, developer retention and future DAO allocations in one measurement frame.
Those are not one class of funded object.
STIP and LTIPP were primarily protocol/user/liquidity incentive programs. Their question is closer to:
incentive
→ induced activity / liquidity
→ post-incentive retention
Stylus Sprint and Developer Tooling grants are much broader. The original Stylus program explicitly contemplated:
- production contracts;
- Rust libraries;
- SDK contributions;
- debugging and development frameworks;
- reference code;
- migration tooling;
- education / DevRel materials.
The completed grants already demonstrate why this matters.
SolDB can create durable value through debugger integration and upstream SDK/docs adoption.
Stylus Toolkit is a CLI/package and local development environment; package usage and downstream developer adoption are meaningful outputs.
Benchmarking Stylus: Common Contract Primitives is a reproducible benchmark/reference artifact. Its value is not equivalent to keeping one production dApp contract busy forever.
A universal rule such as:
more on-chain gas / more deployments
=
more grant impact
will generate false negatives for good infrastructure and false positives for easy-to-manufacture activity.
Minimal repair
Add a frozen ProgramClass / ImpactClass before evaluation.
A workable initial taxonomy is:
INCENTIVE_LIQUIDITY
PRODUCTION_PROTOCOL
DEVELOPER_TOOLING
LIBRARY_OR_SDK
RESEARCH_OR_BENCHMARK
EDUCATION_OR_DEVREL
This taxonomy does not need to be perfect on day one.
It needs one invariant:
a metric is admissible only
if the funded objective class authorizes that metric
For the first Arbitrum version, I would actually narrow the formal mandate to Stylus / Developer Tooling and add STIP/LTIPP adapters later.
That makes the first build smaller and more defensible.
The current claim that one Stylus Oracle can provide the baseline for all future Arbitrum allocations is broader than the current specification supports.
3. Module B currently treats gas as more information than it really contains
The current Sybil Resistance Vector proposes to cross-reference public repository commits with on-chain gas expenditure to filter low-effort activity and identify high-value protocol engineers.
This is the most dangerous anti-gaming boundary in the present specification.
These objects are different:
GitHub contributor
!= funded entity
!= wallet controller
!= deployer
!= gas payer
!= application user
!= downstream developer
A Git commit author field is not controller proof.
A wallet is not a person.
A gas payer may be a relayer, bundler, paymaster or sponsor.
And one controller can create many wallets, many contracts and large amounts of self-generated activity.
There is also a Stylus-specific inversion here.
The original Stylus funding rationale explicitly values efficiency. A successful Stylus implementation may reduce the gas required to produce the same useful work.
So raw gas expenditure cannot safely be treated as a value score.
If gas expenditure is positively scored as a proxy for value, the Oracle can create a perverse incentive:
less efficient implementation
→ higher gas expenditure
→ apparently higher "impact"
That would be the opposite of what Stylus is supposed to reward.
Minimal repair
Keep GitHub and gas, but downgrade both to what they actually prove.
GitHub activity = contribution evidence
on-chain gas = activity evidence
Neither should directly prove:
human identity
independence
quality
economic value
Module B should first resolve an actor/control graph:
FundedEntity
Contributor
Wallet
Deployer
Contract
GasPayer
IndependentUser
DownstreamBuilder
Then classify observed activity into evidence states such as:
TEAM_CONTROLLED
TEST_OR_DEV
SELF_GENERATED
SPONSORED
EXTERNAL_INDEPENDENT
CIRCULAR_OR_COORDINATED
UNKNOWN_CONTROLLER
The exact names can change. The separation cannot.
The real target metric is not “gas spent.”
It is closer to:
qualified independent retained activity
under an objective-specific rule.
I would also remove or narrow the phrase “high-value protocol engineer” from the deterministic layer.
The Oracle can establish qualified activity and integrity evidence.
It should not infer the value of a human being from commits and gas.
4. “Budget-to-outcome linkage” requires a funding lineage, not only an activity lineage
The motivation correctly identifies a missing budget-to-outcome link.
But the current modules do not yet define the budget side of that link.
For grants and incentive programs:
requested
!= approved
!= committed
!= disbursed
!= distributed / used
!= returned
If the Oracle calculates cost-per-builder from an approved budget while part of the money was never paid, the denominator is wrong.
If it treats a transfer to a protocol as final ecosystem expenditure when some of that money is later returned, the denominator is wrong again.
For ARB-denominated historical programs, USD comparisons also depend on a declared valuation convention.
Minimal repair
Bind a small financial ledger to every evaluated object:
funding_source
proposal / grant decision id
asset
approved_amount
committed_amount
finalized_disbursements
payment_tx_set
milestone_id
recipient_address
returned_amount
unspent_amount
valuation_method_version
Then define the cost object once.
For example:
NetDisbursedCost
=
finalized disbursements
-
finalized returns
If a USD conversion is displayed, retain both:
native asset amount
+
declared historical valuation rule
Do not overwrite the historical native amount with a mutable current USD value.
For incentive programs, an additional distinction may be required:
DAO → protocol
!=
protocol → final incentive recipients
The Oracle should know which financial transition it is claiming to measure.
Without this layer, “budget-to-outcome linkage” is still only outcome telemetry with a budget number attached to it.
5. Retention, cost efficiency, ROI and integrity are four different claims
The present RFC uses “ROI Matrix”, “cost-per-builder”, “retention” and Sybil filtering very close together.
They should be separated before implementation.
These are different claims:
Delivery:
Was the funded obligation delivered?
Retention:
Did the qualified funded object remain active after t?
Adoption:
Did independent external actors use or depend on it?
Cost efficiency:
How much net funding was spent per qualified retained unit?
Causal impact:
How much of the observed outcome is attributable to the grant?
Integrity:
Is there evidence that reported activity was manipulated or misrepresented?
A valid descriptive metric can be:
ObservedCostPerRetainedBuilder
=
NetDisbursedCost
/
QualifiedRetainedBuilderCount
But that is not automatically “ROI.”
If ROI is intended to mean grant-attributable impact, then a causal boundary is required:
observed post-grant activity
!=
incremental activity caused by the grant
The team may have continued building without the grant.
A market-wide Stylus expansion may have created activity.
Another incentive may have caused the retention.
A good project can underperform because of an external regime change.
Minimal repair
Attach an attribution class to every impact claim:
DESCRIPTIVE_ONLY
BASELINE_ADJUSTED
CONTROL_MATCHED
CAUSAL_SUPPORTED
Do not promote a claim above the evidence class that supports it.
If the first version has no defensible counterfactual, call the output:
Impact Efficiency Matrix
or:
Observed Grant Outcome Matrix
and reserve ROI for fields whose return and attribution are actually defined.
This does not weaken the Oracle.
It makes its conclusions much harder to attack.
One more separation is essential for anti-fraud use
Low performance must not automatically become an integrity accusation:
LOW_RETENTION
!=
FRAUD
and:
HIGH_ACTIVITY
!=
CLEAN
Keep two independent states:
PerformanceState
IntegrityState
A future funding process may decide how to use both.
The Oracle should not silently convert one into the other.
6. Restore the five evidence states as real state, not dashboard labels
The five-state boundary accepted in the August 24 discussion is currently absent from the public Specifications section.
That should be restored explicitly.
Every material qualification decision should preserve at least:
status
reason_code
evidence_root
methodology_version
input_snapshot_hash
decision_id
predecessor_decision_id
effective_time
review_status
with:
ACCEPTED
EXCLUDED_WITH_REASON
DISPUTED
UNRESOLVED
SUPERSEDED
Three fail-closed rules matter here
A. Missing evidence is not zero
insufficient evidence
→ UNRESOLVED
not:
insufficient evidence
→ inactive / failed / excluded
B. Exclusions cannot disappear
If 40 objects are excluded, the dashboard should expose that fact and the reasons.
Otherwise the denominator itself becomes gameable.
C. Methodology changes cannot silently rewrite history
If methodology v2 changes a threshold, the original v1 result should remain addressable.
The new result can supersede it:
Decision_v1
→ SUPERSEDED_BY
Decision_v2
but v1 should not vanish.
Every important dashboard number should therefore travel with something like:
value
eligible_n
accepted_n
excluded_n
unresolved_n
disputed_n
methodology_version
observation_epoch
evidence_snapshot
That makes uncertainty visible instead of laundering it into a clean percentage.
7. Dune should be a read model, not the trust root
Version-controlling SQL through public GitHub PRs is useful.
It is not sufficient by itself for a zero-trust claim.
A query can be perfectly versioned while its inputs remain mutable.
The evaluation needs to bind:
cohort manifest
grant metadata snapshot
repository / release identifiers
SQL commit
methodology hash
chain id
finalized block range
observation timestamp / epoch
source-data version where applicable
Then a third party should be able to reproduce the classification independently from the published evidence package.
The architecture should therefore be:
Canonical evidence + methodology
→ reproducible computation
→ Dune / dashboard presentation
not:
Dune result
→ canonical truth
I would also soften zero-trust maintenance to trust-minimized maintenance until independent replay is actually part of the deliverable.
8. Governance needs to separate four different kinds of change
The 3-of-5 PR gate is useful, but “parameter update” currently covers too many semantic operations.
At minimum distinguish:
CODE_FIX
DATA_CORRECTION
METHODOLOGY_CHANGE
INDIVIDUAL_RECORD_APPEAL
They should not all have the same effect.
A code fix may repair an implementation bug.
A data correction may repair one record.
A methodology change changes what the Oracle means.
A record appeal challenges one classification under an already frozen methodology.
Two important rules
First:
methodology selected after observing the result
must not silently become:
methodology that was always in force
Freeze an effective methodology version for each evaluation epoch.
Second, I would not hard-code the governance authority to the name of one interface such as Snapshot.
Bind the system to an ArbitrumDAO governance decision / decision ID under the then-current DAO procedure.
Interfaces can change.
Authority should survive the UI.
DAO-level methodology disputes and individual record disputes should also be separate paths. Requiring a DAO-wide vote for every disputed wallet, contributor or contract will not scale.
9. The fastest way to make this proposal-ready is a bounded pre-build pilot
I would not begin the $20,000 SQL milestone by writing the broad indexer.
Add one small acceptance gate before Milestone 1:
Milestone 0 — Qualification Specification & Adversarial Pilot
Take a small set of already-public, completed grants representing different objective classes.
A useful initial set is:
- Stylus Mobile Native Authentication
- SolDB — CLI debugger and simulator for Solidity and Stylus
- Stylus Toolkit
- Benchmarking Stylus: Common Contract Primitives
- ZeroStyl — Privacy Toolkit
For each one, freeze:
funded obligation
funding lineage
objective class
source / artifact lineage
qualification state
controller / independence state
retention anchor
admissible impact metrics
attribution class
final reconciliation
Do this manually first.
That produces the gold schema the SQL engine must later reproduce.
The pre-build adversarial tests should include at least these cases
Test 1 — Controller multiplication
one controller
→ 50 wallets
→ 50 deployments
→ self-funded gas
Required result:
not 50 independent builders
Test 2 — Artifact laundering
unrelated fork / clone
→ similar interface or bytecode family
→ active deployment
Required result:
does not count as funded-artifact retention
without proven lineage
Test 3 — Tooling false negative
debugger / SDK / package
→ upstream integration or real downstream usage
→ little persistent gas footprint
Required result:
can succeed under its objective class
Test 4 — Gas-payer misbinding
real user
→ relayer / bundler / paymaster pays gas
Required result:
gas payer is not promoted to user or builder identity
Test 5 — Upgrade continuity
qualified proxy / implementation
→ authorized upgrade
→ new code hash
Required result:
lineage continues only through a proven upgrade edge
Test 6 — Missing provenance
activity exists
but source/build/controller evidence is insufficient
Required result:
UNRESOLVED
not zero, failure or silent exclusion.
Test 7 — Funding correction
approved amount > actual disbursement
or unused funds are returned
Required result:
cost metric uses the declared net-cost rule
Test 8 — Methodology revision
v1 classifies object
v2 changes threshold
Required result:
v1 remains auditable
v2 creates a new decision
v1 is linked as SUPERSEDED where appropriate
If the schema survives those cases, the Dune implementation has a stable target.
If it does not, more SQL will only make the ambiguity faster.
10. Exact changes I would make to the current RFC
This does not require rewriting the proposal from scratch.
In Abstract
Narrow:
"historical ROI ... baseline for all future Arbitrum DAO allocations"
to an objective-qualified impact / cost-efficiency claim for the first supported program classes.
Expand scope later through explicit adapters.
Before current Module A
Insert:
Precondition Layer — Grant, Funding & Artifact Qualification
This restores the provenance/evidence gate accepted on August 24.
Rewrite Module A
Current concept:
cohort wallet
→ deployment
→ strict 30/90/180 post-grant
Required concept:
qualified object
→ objective-specific finalized t0
→ qualified 30/90/180 observation windows
→ retention result + coverage state
Rewrite Module B
Do not compute engineer value from:
GitHub commits × gas expenditure
Use:
identity / controller resolution
+
independence classification
+
objective-admissible activity
+
negative controls
GitHub and gas remain evidence inputs, not final semantic outputs.
Expand Module C
Add:
methodology versioning
effective epochs
non-retroactive history
record-level appeal
SUPERSEDED lineage
Keep DAO sovereignty over methodology, but do not make one DAO-wide voting path the adjudicator of every individual data dispute.
Add an output contract
Every published metric should expose:
what object was measured
what evidence class supports it
what methodology version produced it
what denominator was used
how much of the cohort is unresolved / disputed / excluded
whether the result is descriptive or causally attributed
Before current Milestone 1
Add the bounded five-grant qualification pilot above as a pass/fail gate.
Then the SQL engine has an exact specification to implement.
11. Two small public-specification inconsistencies are also worth fixing
The closing note says that the foundational invariants and evidence-binding requirements are mapped in “Section 3.”
In the public RFC as currently posted, those invariants are not actually present in the Specifications section, and the five-state model/provenance gate discussed and accepted on August 24 is absent.
If there is a separate Section 3 draft, link it.
If not, restore that material directly into the RFC so delegates are reviewing the same architecture that the implementation will use.
There is also an open-source timing ambiguity:
- Module C says SQL logic is version-controlled through public GitHub PRs.
- The Public Good Clause says the codebase and SQL schema will be open-sourced post-60-day.
Clarify which components are public from day one and which are released at the end.
For governance infrastructure, I would strongly prefer the active methodology and SQL to be public from the first evaluation epoch.
Bottom line
I agree with the core direction.
The Oracle can become useful infrastructure for making grant gaming harder, but the anti-gaming value will come from proof boundaries, not from adding more counters.
The first invariant should remain:
Local measurement does not authorize semantic promotion.
So:
deployment
!= funded artifact
gas
!= independent adoption
GitHub commit
!= funded builder identity
retention
!= causal grant impact
low impact
!= fraud
high activity
!= integrity
deterministic SQL
!= correct semantics
The good news is that the repair is bounded.
You do not need to redesign the whole Oracle.
Restore the qualification gate you already accepted, type the funded objects, separate performance from integrity, bind the funding ledger, make the evidence states/version history first-class, and run the five-object adversarial pilot.
After that, the SQL/indexer/dashboard layer becomes much safer to build.
@stonecoldpat — moving this to Early Idea Discussion is reasonable at the current stage. I think the shortest path from here to a proposal that can be evaluated against the 50k ask is to make the pre-Milestone-1 qualification specification a concrete acceptance artifact rather than starting with the dashboard.
The first frozen qualification matrix can be mapped directly from the public grant reports and repositories; no private access is required.
Public references