Arbitrum Audit Program: Transparency Report #3

Operational Period: February 01, 2026 – April 30, 2026

The DAO-approved Arbitrum Audit Program (AAP) completed its third operational quarter during the period from February 01, 2026, to April 30, 2026 (“Q3”). Launched on August 01, 2025, the program runs an open application process for one year to support teams seeking audit subsidies to improve the security and reliability of their projects.

Q3 was the program’s busiest quarter since launch in terms of application volume. The committee reviewed 108 applications, representing a 56% increase compared to Q2 and a 33% increase compared to Q1. Despite the rise in submissions, evaluation standards remained consistent, with 12 projects initially approved, reflecting an 11% acceptance rate. However, the quarter also saw a notable increase in post-approval withdrawals, reducing the number of projects that ultimately proceeded through the program.

Across these approved audits, 14 have been completed and security firms identified 297 vulnerabilities, including 8 classified as critical, while reviewing a total of 21,882 lines of code. These findings underscore the importance of rigorous security review in strengthening the resilience and reliability of applications building within the Arbitrum ecosystem.

While application volume increased significantly during the quarter, the proportion of mature, audit-ready teams did not expand at the same pace. The committee absorbed this increased demand without lowering evaluation standards, maintaining a focus on supporting projects with strong technical readiness, ecosystem alignment, and long-term potential.

Key highlights

  1. 108 applications received during Q3, with DeFi remaining the most prominent category.

  2. 12 projects were initially approved through committee evaluation, representing an 11% approval rate. Following post-approval withdrawals, 7 projects ultimately continued through the program.

  3. Across 14 completed audits, 297 vulnerabilities were identified (including 8 classified as critical and 31 as high), and 21,882 lines of code were reviewed.

  4. Approximately $1.12 million has been committed for approved audits since launch, with total audit commitments expected to reach approximately $2 million, including audits currently in progress or pending execution.

Application Pipeline Analysis

Highest Application Volume Since Launch

The program received 108 applications during Q3, marking the highest quarterly intake since launch. For comparison, the program received 81 applications in Q1 and 69 applications in Q2. The increase in submissions reflects the cumulative impact of ongoing ecosystem outreach and marketing efforts, but more notably, the growing role of referrals in driving high-quality deal flow. Of the 7 projects ultimately onboarded, 5 originated through referrals from audit firms, ecosystem partners, or previously supported teams. Projects that successfully completed audits helped expand awareness, effectively becoming advocates for the initiative.

Despite the substantial increase in application volume, the committee maintained a consistent review framework and evaluation standard throughout the quarter. As a result, approval rates declined relative to prior quarters, falling to 11%, compared to 13% in Q2 and 20% in Q1. While 12 projects initially received approval, 5 later withdrew from the program, resulting in 7 projects being fully onboarded. One of the approved engagements represented an extension of a previously supported audit.

Post-Approval Withdrawals

Q3 saw a higher-than-usual number of post-approval withdrawals, where projects exited the program after initially receiving committee approval. These withdrawals were primarily driven by technical delays and misalignment with certain program requirements, particularly around ecosystem exclusivity expectations for supported deployments. However, introducing greater flexibility around exclusivity requirements enabled the program to invite applications from teams that would otherwise have been ineligible for support.

The withdrawals broadly fell into two categories:

  • Some teams secured integration or partnership opportunities that required deployment environments outside of Arbitrum, leading them to voluntarily withdraw from the program.

  • Other teams decided to significantly rework portions of their codebase or feature set after approval and were unable to provide a reliable timeline for delivering an audit-ready implementation. These teams have been encouraged to reapply once development stabilises and a revised codebase is ready for review.

The committee is actively monitoring this trend and evaluating process improvements that may help reduce post-approval attrition in future while preserving flexibility for high-quality teams navigating evolving product roadmaps.

Application Quality

The primary driver behind high Q3 rejections was overall application quality. A large portion of applicants had limited operational and technical maturity, with many submissions coming from solo founders or early-stage teams whose products, value propositions, or codebases were not yet sufficiently developed for a formal security audit process.

In many cases, these teams would likely benefit more from earlier-stage ecosystem support initiatives, such as Open House and Arbitrum Mentorship program, before applying to the Arbitrum Audit Program. The teams were consequently guided towards these programs.

The second major factor behind rejections was insufficient alignment with the program’s ecosystem requirements and exclusivity expectations. Several applicants were unable to clearly articulate a credible Arbitrum-focused growth strategy, while others were unwilling or unable to commit to the deployment and ecosystem alignment conditions associated with the program.

Compared to previous quarters, the underlying rejection patterns remained broadly consistent. However, the increase in overall application volume was accompanied by a proportional increase in lower-maturity submissions. This significantly expanded the committee’s review workload without resulting in a corresponding increase in high-quality, audit-ready projects progressing through onboarding.

Applicant Composition

The applicant pool reflected a healthy mix of new teams discovering the Arbitrum ecosystem alongside more established ecosystem participants seeking security support.

DeFi remained the dominant category, accounting for 54% of all applications (58), followed by Infrastructure (30) and AI (7). Among the projects that successfully completed onboarding without cancellation or withdrawal, DeFi remained the most represented category. Three of the approved projects – Variational, Nashpoint, and ARU Reserve – were DeFi-focused, while Superset represented the Infrastructure category.

The composition of approved projects would have been notably more diverse had all initially approved teams proceeded through the program. The withdrawn or paused projects included teams across gaming, infrastructure, AI, and DeFi, indicating that interest in the program continues to expand beyond DeFi.

Audits Completed & Security Findings

Across approved audits, 14 have been completed since launch, of which five are in production. The following projects have completed audits through the Arbitrum Audit Program:

  1. Bleap
  2. Triumph Games
  3. Cybro
  4. Kandle Finance
  5. idOS
  6. Footium
  7. Tezoro
  8. Kleros
  9. Nashpoint 1
  10. Nashpoint 2 (extension of initial audit)
  11. Stormbit Finance
  12. Capx AI 1
  13. Capx AI 2 (extension of initial audit)
  14. Stimpak (Duels)

Across all approved audits, the program has committed a total of $1,225,120.
Across all 14 completed audits, the average audit cost is $46,363. The average audit cost per line of code was $45/LoC. However, for audits covering codebases larger than 3,000 LoC, the average cost falls to approximately $15 per LoC. This shows the scale of audit pricing and provides important context when comparing the cost of larger audits.

While several cancelled or withdrawn audits reduced near-term expenditure during Q3, audits currently in progress, paused, or pending execution are expected to bring the program’s total committed capital closer to approximately $2 million over time.

Across all 14 completed audits, auditors identified 297 vulnerabilities, including 8 classified as critical severity issues, with the potential to pose significant risk if they were not detected as part of the AAP. A total of 21,882 lines of code were audited under the completed audits.

Budget Deployed

As of April 30, 2026, the program has committed $1,225,120, representing approximately 12% of the total $10 million annual budget. While budget deployment remains conservative relative to the total allocation, this reflects a deliberate decision by the committee to maintain consistent selection standards throughout the program’s lifecycle.

As noted in the previous transparency report, participating teams have consistently requested higher audit coverage, often up to 100% of audit costs, in exchange for accepting the program’s exclusivity requirements. In most approved cases, the program accommodated these requests. As a result, while the total number of financed teams remains relatively modest compared to the available budget, supported teams have generally received deeper financial support per engagement.

Audits currently paused, pending activation, or under negotiation are expected to bring cumulative commitments to approximately $2 million once activated, particularly as several pipeline projects involve larger and more complex codebases requiring more extensive review scopes. A portion of the remaining budget is also expected to be allocated toward the planned AI service provider pilot. While pricing discussions with prospective providers are still ongoing, the expected allocation for this initiative remains negligible relative to the overall program budget.

Auditor Participation & Expansion

During Q3, the program’s roster expanded to 13 approved audit providers. Trail of Bits and Blackthorn (Sherlock) were added as two new auditors following the completion of the program’s due diligence and onboarding process.

Both firms had previously participated in the initial auditor review and whitelisting process during the program’s launch phase but were not onboarded at that time. Their addition during Q3 reflects continued demand from applicants and broader ecosystem interest in working with these firms.

The committee has consistently observed that applicants tend to show strong preferences, with many teams explicitly requesting to work with specific providers based on reputation, prior working relationships, or technical specialisation. The addition of new high-quality firms, therefore, helps expand auditor availability while aligning with applicant demand.

The 13 approved auditors participating in the program during Q3 were:

  1. OpenZeppelin
  2. Certora
  3. Nethermind
  4. Ackee Blockchain Security
  5. Oak Security
  6. Hexens
  7. Decurity
  8. Pashov Audit Group
  9. OXORIO
  10. Cyfrin
  11. Guardian
  12. Trail of Bits
  13. Blackthorn (Sherlock)

Revisiting the Exclusivity Requirement

Following discussions raised in previous quarters, a formal governance proposal was submitted to the DAO to revisit the program’s exclusivity requirement. The proposal introduced a shift from a strict mandatory exclusivity condition toward a more flexible ecosystem alignment framework.

Under the revised framework, Arbitrum exclusivity continues to remain the preferred standard for supported projects. However, the committee now has limited discretion to grant exemptions in cases where projects demonstrate strong strategic alignment with the Arbitrum ecosystem despite operating in a broader multi-chain context.

The updated framework has since been approved and applied selectively to a small number of projects. During Q3, three projects received exemptions under the revised policy, including Superset, a cross-chain stablecoin and foreign exchange infrastructure project that selected Arbitrum as its primary hub chain and could drive substantial transaction volume and network activity.

Audit Completion Timeline

The AAP approaches its scheduled completion on 31 July 2026. Applications will remain open until 31 July.

Audits approved prior to the application deadline will continue to be supported through completion, even if they extend beyond the program’s close date. To accommodate this, we anticipate an additional two-month wind-down period following the application deadline to process and finalise all pending audits before the program can be fully sunset.

A final transparency report will be published once all approved audits have been completed and all program commitments have been fulfilled.

3 Likes

@Arbitrum, I reviewed the original DAO-approved Audit Program specification, Transparency Reports #1#3, the auditor selection and launch notices, the April program-improvement proposal, and the July GRC update confirming that five completed audits are now being benchmarked against five AI security tools.

The program has clearly established a functioning intake, selection, procurement, and audit-delivery mechanism.

The remaining issue is acceptance identity.

At present, “audit completed” principally establishes that an audit report was delivered, accepted by the project, reviewed by the Foundation for report quality, and became eligible for payment.

That is a valid procurement milestone.

It is not yet the same object as:

  • findings remediated;
  • remediation independently retested;
  • residual risk explicitly accepted;
  • corrected code frozen;
  • deployed bytecode bound to the corrected code;
  • post-deployment security outcome observed.

This distinction matters now because the program is approaching redesign, five of the fourteen completed engagements are reported as being in production, and the AI Security Pilot is already comparing five audits with five AI tools.

1. Separate report completion from security acceptance

Every engagement should move through explicit states:

SCOPE_FROZEN
AUDIT_EXECUTED
AUDIT_REPORT_ACCEPTED
FINDINGS_ADJUDICATED
REMEDIATION_SUBMITTED
RETEST_COMPLETE
RESIDUAL_RISK_ACCEPTED
DEPLOYMENT_BOUND
POST_DEPLOYMENT_OBSERVED

These states should not be collapsed into COMPLETED.

The existing payment rule can remain tied to AUDIT_REPORT_ACCEPTED if that is the contractual model.

However, program-level claims about security impact should be derived from the later states.

For example:

  • an accepted report proves that a professional review was delivered;
  • a passed retest proves that a specific correction closed a specific finding;
  • deployment binding proves that the accepted corrected artifact is the artifact operating on-chain;
  • post-deployment observation provides evidence about escaped defects and operational performance.

2. Create one privacy-preserving acceptance receipt per engagement

The program does not need to disclose confidential source code, private findings, individual prices, or sensitive project details.

It can still publish a detached receipt containing:

  • engagement ID;
  • project identity;
  • selected auditor;
  • audit type: initial / extension / retest;
  • exact scope commit or source-tree hash;
  • dependency-lock hash;
  • compiler and build identity;
  • scope start and end dates;
  • report hash;
  • normalized finding counts;
  • remediation commit or source-tree hash;
  • retest report hash;
  • open and accepted residual-risk counts;
  • target chain;
  • deployed contract addresses;
  • deployed runtime-code hashes;
  • deviations from the audited build;
  • final acceptance state.

This closes the chain:

frozen scope
→ audit report
→ findings
→ corrections
→ retest
→ deployment

Without that chain, “14 audits completed” measures report production, while “five are in production” does not establish that the deployed artifacts are the corrected artifacts reviewed by the auditors.

3. Normalize findings before aggregating them

The reported 297 vulnerabilities are useful evidence of audit activity.

They are not yet a comparable program-level security metric unless all providers are mapped into one normalization contract.

For every finding, the program should distinguish:

  • provider severity;
  • normalized program severity;
  • unique finding identity;
  • duplicate or overlapping finding;
  • confirmed;
  • disputed;
  • accepted risk;
  • fixed;
  • partially fixed;
  • retest passed;
  • retest failed;
  • not retested;
  • affected deployed artifact;
  • escaped into production.

This is particularly important where:

  • multiple providers use different severity taxonomies;
  • an extension rechecks overlapping code;
  • the same root cause produces several symptoms;
  • one report counts informational or gas issues and another does not.

A useful public aggregate would be:

valid unique findings
fixed
retested
deployed with fix
remaining residual risks

That sequence measures risk reduction more directly than raw finding count.

4. Reconcile the current reporting identities

Several published values should be reconciled before the final report.

Approval-rate chronology

Transparency Report #3 states that Q3’s 11% approval rate compares with 13% in Q2 and 20% in Q1.

Transparency Report #2 records:

  • Q2: 14 / 69, approximately 20%;
  • Q1: initially 11 / 81, approximately 13%;
  • Q1 after two later approvals: approximately 16%.

The Q1 and Q2 values in Report #3 therefore appear to have been reversed, and the revised Q1 denominator state is not preserved.

Commitment total

The highlights state approximately $1.12 million committed.

The detailed sections state $1,225,120.

Those values differ by $105,120. The latter would normally round to approximately $1.23 million.

Completed-audit denominator

The text reports fourteen completed audits.

The published cost-per-LoC chart appears to contain three buckets with counts:

  • 7;
  • 5;

Those counts total fifteen. Their weighted mean also reproduces the reported average of approximately $45/LoC, suggesting that fifteen observations were used in that calculation.

The additional observation, extension treatment, or chart denominator should be identified.

Provider lifecycle

Trail of Bits was included in the July 2025 final accepted-firm list and in the August launch notice describing twelve onboarded firms.

Transparency Report #1 subsequently listed eleven active/approved firms without Trail of Bits.

Transparency Report #3 then described Trail of Bits as newly added during Q3 and stated that it had previously been whitelisted but not onboarded.

These statements may all refer to different valid lifecycle states, but those states need canonical definitions:

APPLIED
PASSED_DUE_DILIGENCE
ACCEPTED
WHITELISTED
CONTRACTUALLY_ONBOARDED
ACTIVE
SUSPENDED / RETIRED

Auditor attribution

The original approved reporting specification states that completed audits should be disclosed together with the selected auditor.

Report #3 lists the completed projects and the overall provider roster, but does not map each completed engagement to its selected auditor.

A public engagement-to-provider mapping would satisfy the original reporting commitment without disclosing confidential pricing.

5. Separate financial states and publish the calculation contract

The current section is titled Budget Deployed, but the reported value is funds committed.

The final reporting model should separately record:

  • approved;
  • quoted;
  • contracted;
  • committed;
  • invoiced;
  • paid;
  • refunded;
  • cancelled;
  • reserved;
  • remaining treasury balance.

The cost statistics should also identify:

  • whether “audit cost” means total invoice or DAO subsidy;
  • whether the average is weighted or unweighted;
  • whether extensions are separate observations;
  • whether overlapping code is counted again;
  • the exact LoC definition;
  • whether generated code, tests, interfaces, dependencies, comments, and unchanged inherited code are included;
  • the denominator used for every chart.

LoC can be a useful scope descriptor, but it is not independently comparable without a frozen counting policy.

6. Preregister the AI benchmark before interpreting its outputs

The July update states that five full audits are being benchmarked against five AI security tools.

This is exactly the point at which the benchmark contract should be frozen.

Every tool should receive:

  • the same exact source-tree hash;
  • the same dependency closure;
  • the same declared scope;
  • the same compiler and environment assumptions;
  • the same time and resource limits;
  • the same permitted external context;
  • no access to the reference audit report or hidden answer set.

Each run should bind:

  • provider identity;
  • tool and model version;
  • configuration;
  • prompt or orchestration identity;
  • execution timestamp;
  • context supplied;
  • output manifest;
  • human review time;
  • cost.

The reference audit report must not automatically become complete ground truth.

A human audit can also omit a real defect.

The benchmark therefore needs an independent adjudication layer:

  1. freeze the valid findings from the reference audit;
  2. normalize and deduplicate them;
  3. review every AI-only finding independently;
  4. classify it as valid, invalid, duplicate, out of scope, or unresolved;
  5. record reference findings missed by each tool;
  6. preserve findings that neither the original auditor nor the AI detected but that are later discovered.

Required metrics should include:

  • valid-finding precision;
  • recall by normalized severity;
  • Critical/High false-negative count;
  • duplicate burden;
  • invalid-finding burden;
  • reviewer time required;
  • reproducibility across repeated runs;
  • time to first valid finding;
  • total cost;
  • remediation usefulness;
  • unsupported confidence or false-assurance events.

7. Keep the retrospective benchmark separate from the intended service workflow

The intended AI service is for early-stage teams that are not yet ready for a full audit.

The current benchmark uses completed full audits.

Those are different populations.

The pilot should therefore contain two separate studies.

Retrospective paired benchmark

Purpose:

Determine what each tool detects on the same frozen code that later received a professional audit.

Prospective readiness study

Purpose:

Determine whether an early-stage team that receives an AI-assisted scan:

  • resolves valid issues;
  • reaches audit readiness faster;
  • receives fewer valid findings during the later human audit;
  • reduces audit duration or cost;
  • avoids deploying before it is ready.

The prospective outcome should not be:

AI scan completed → code is secure

It should be:

PRE_AUDIT_READINESS: READY / PARTIAL / NOT READY / BLOCKED

with explicit evidence boundaries and unresolved areas.

8. Make benchmark adjudication independent

Where a provider supplies both:

  • an AI tool;
  • and professional human audits,

the provider should not be the sole judge of its own benchmark performance.

The benchmark should record conflicts of interest and use an independent adjudicator for:

  • finding validity;
  • severity normalization;
  • duplicate resolution;
  • out-of-scope decisions;
  • false-negative classification;
  • final pass/fail interpretation.

The same independence principle should apply to final audit acceptance: the primary auditor may verify its own fixes, but a high-consequence deployment benefits from a separate final reviewer who did not produce the original report.

9. Separate program activity, security outcome, and ecosystem outcome

The current reports provide strong activity metrics:

  • applications;
  • approvals;
  • audit engagements;
  • findings;
  • LoC;
  • commitments;
  • provider participation.

The final report should add two other layers.

Security outcome

  • findings fixed;
  • findings retested;
  • residual risks accepted;
  • audited builds deployed;
  • deviations from audited builds;
  • post-launch incidents;
  • escaped Critical/High defects;
  • time from report to remediation closure.

Ecosystem outcome

  • projects launched;
  • projects still operating at 2, 4, and 6 months;
  • Arbitrum deployment retained;
  • usage, TVL, fees, integrations, and users;
  • audit subsidy as a causal factor in launch;
  • projects that would otherwise have deployed elsewhere or not deployed;
  • subsidy spent on projects that never reached production.

This closes the difference between:

the DAO funded audit activity

and:

the DAO produced measurable security and ecosystem value.

Recommended acceptance architecture for the next program

I would structure the next phase around four independent gates:

Gate 1 - Audit readiness

Exact scope, stable code, build reproducibility, documentation, dependency closure, threat model, and deployment assumptions are frozen.

Gate 2 - Audit delivery

The auditor completes the agreed scope and the report passes contractual quality review.

Gate 3 - Security closure

Findings are adjudicated, remediations are mapped, fixes are retested, and residual risks are explicitly accepted.

Gate 4 - Deployment binding

The deployed runtime artifacts are cryptographically bound to the accepted corrected build, and any later deviation reopens acceptance.

The highest-value next public artifact would be:

  1. one privacy-preserving acceptance receipt for a completed project already in production; and

  2. one preregistered methodology for the five-audit / five-tool AI benchmark.

Those two objects would allow the DAO to evaluate the next version of the Audit Program from reproducible security outcomes rather than from report counts alone.

The program has already built the difficult operational layer.

The next step is to make the transition from audit delivery to verified risk reduction explicit, measurable, and independently reviewable.

1 Like