Clean Weights: alignment-audited open models for regulated enterprises
The problem
Chinese open-weight models now power roughly half of OpenRouter traffic and are the cheapest capable base models available, but the alignment travels in the weights: Booz Allen found Qwen produced 130% more security vulnerabilities when told a project was for the US government, CrowdStrike found DeepSeek wrote 50% more vulnerable code when the task was framed against Chinese interests, and Harmonic Security found 1 in 12 employees already use Chinese GenAI tools at work, leaking code and keys. Enterprises want the capability and the cost, not the hidden alignment.
The idea
Why now
Hirundo's Oct 5 Westernized Qwen release proved weight-level political unlearning works in practice (89.8% to 2.8% CCP-aligned answers with capabilities held within 0.72 points) and the lab plans to publish its CCPC-500 benchmark; OpenRouter's leaderboard shows 8 of the top 10 models are Chinese, so demand for a trusted path to those weights is at its peak; and the September NSA/CISA/FBI distillation advisory has enterprise security teams scrutinizing Chinese models harder than ever.
What it combines
Hirundo's behavioral weight-editing (models/open-weights) plus Jev-Omni typed judgments (models/decision): unlearning removes the behavior, and decision classifiers turn auditing into an automated pass/fail line. The combination makes weight certification repeatable at scale, which neither a one-off lab demo nor a generic API auditor can do.
MVP
Weekend scope: audit-only pipeline for one model family (Qwen3.8) - download the weights, run a public political benchmark plus capability benchmarks, publish the attestation page. Deliberately skip the unlearning step in v1, multi-family support, and hosting the weights.
Distribution
B2B2C: sell the certification to inference platforms (Together, Fireworks, Nebius, OpenRouter) so they can badge hosted models as audited - the platforms already own the developers. Second motion: direct to regulated enterprises and government-adjacent buyers who must demonstrate due diligence.
Why it wins
Hirundo sells its own unlearning service and no independent third party issues comparable attestations; existing AI auditors (Holistic AI, Robust Intelligence) audit API behavior, not weight-level alignment. Nobody offers a certified-weights marketplace with signed attestation cards.
Risks
Benchmark gaming: providers optimize for the published tests. The MVP de-risks by keeping part of the eval prompt set private and rotating it, and by publishing the methodology so gaming is visible.
Build it with
- Hirundo Westernized Qwen: CCP political alignment unlearned from Qwen weightsThe behavioral weight-editing method to replicate for the unlearning pass.
- Jev-Omni — 12B multimodal decision classifierTyped decision judgments to automate pass/fail auditing across thousands of prompts.
- DoubtBench – does Jev know when humans disagree?Validate where the judgment layer is uncertain when humans disagree, so borderline audits get human review.
Repo to start from
clean-weights-audit - a pipeline that downloads an open-weight release, runs political-alignment and security-behavior differential audits, and publishes a signed attestation card.
Evidence
- Hirundo businesswire release
- Capital Digest on Qwen censorship and security-lab findings
- Harmonic Security on Chinese GenAI workplace usage
- China AI Weekly on the NSA/CISA/FBI distillation advisory
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.