AWAI · AW1
Model card · System card · updated 05/10/2026

AW1, tách đúng
model và engine.

AW1 là một hệ thống production gồm model weights, inference engine và agent/evidence engine. Ba lớp có danh tính và mức bằng chứng độc lập; candidate không được trình bày như production.

English summary. AW1 combines the sieutocviet/aw1 model identity, the AWLLM inference runtime, and a receipt-bound research, rendering and verification system. Unpublished fields remain UNKNOWN; this page is not a claim of frontier-model parity.

Candidate model
AW độc lập
Đang kiểm chứng · chưa promote
Inference engine
AWLLM native
PASS_SCOPED · PRODUCTION

V2 candidate evidence is historical (16/08/2026); current runtime readback: awllm_native_minimal.

System engine
VERA + VIF-E + SEVF
Architecture canon · capability receipts vary

AEO routes work; VERA/VWR owns canonical workstreams; SEVF with AC/ACVL gates evidence; VIF-E/VADR binds evidence to answer lifecycle; WME/ELE credits verified outcomes.

1. Candidate AW — cập nhật 05/10/2026

QUALITY WITHHELD · không promote. Candidate aw1-answerability-entity-ab-20261005-r1 kiểm tra scorer answerability trên QA tổng hợp song ngữ; reader và raw ranking giữ cố định. Đây là thí nghiệm candidate, không phải đánh giá chất lượng production.

Train fitting

Tại checkpoint cold được chọn: control C đạt 8.000/8.000 mỗi ngôn ngữ; reassignment R đạt 9.600/9.600 mỗi ngôn ngữ. Chỉ số này đo quyết định answerability trên dữ liệu fitting, không phải độ chính xác câu trả lời hay held-out.

Development — evidence-group transfer

Hai nhánh C/R có cùng số đếm tại operating point được chọn: EN raw exact 27/40, gated exact 13/40, no-answer đúng 10/20; VI raw exact và gated exact 36/40, no-answer đúng 20/20. Question pairs 13/32; position pairs 19/32. Gate dev FAIL_OR_NOT_RUN; final NOT_RUN_DEV_FAIL; canary NOT_RUN. Không ghi nhận credit cải thiện do reassignment.

Accounting và production

600 cập nhật scorer mỗi nhánh, tổng 1.200 cập nhật answerability; LM updates = 0, reader/start/end updates = 0. Mỗi nhánh có 19.200 binary-label presentations, 16.000 logical rows và 2.000 evidence groups. Candidate chưa chuyển vào production; receipt này không chứng minh chất lượng production.

Điểm chặn: train fitting đã hoàn hảo nhưng không chuyển thành cải thiện answerability trên evidence groups EN mới. QUALITY WITHHELD; không promotion.

Receipt entity A/B 05/10/2026 · SHA256: dff08f355829f4c5496a3875461e74b3d307457cf07967173fdf9497ad795008

Receipt residual A/B trước đó (05/10/2026) có VI gated 35/40, EN gated 13/40, EN no-answer 10/20, question pairs 12/32. Hai receipt dùng packet DEV khác nhau; không coi 35→36/40 là cải thiện đối chứng.

Receipt residual · SHA256: 89d0b2c97bdc8efa79d0b984fc401e4b4cd62b1941ef1b55f616842de86d93ca

Chỉ số lịch sử được công bố ngày 30/09/2026: corpus 1.000.565.587 ký tự từ 6 nhóm nguồn, candidate khoảng 9,3 triệu tham số, 200 bước pilot. Đây là snapshot lịch sử, không cộng 1.200 scorer updates vào LM updates.

Snapshot có numerator, denominator và scope

2. Historical context contract

65,536 native tokens and 64,884 held-out prompt tokens were published on 03/09/2026; no new long-context test was run for this update.

Native production context
65,536 tokens · VERIFIED; standalone merged checkpoint in production
Live production held-out
64,884 prompt tokens · VERIFIED 1/1; dependency after token 32K
Historical reliable effective actual
47,967 tokens · VERIFIED 6/6; retained for lineage
Live production accepted lower bound
64,884 prompt tokens · VERIFIED 1/1; not a claim that every 64K task is solved
Economic context
15,965 tokens · VERIFIED 6/6
Native 64K production verification
13/13 checks PASS · independent_readback_receipt · 0965ef1b…
Observed production request
10,750 tokens · PASS_SCOPED 6/6; not a native limit

primary receipt: RuntimeContextFrontierVNext_20260816T224917Z · sha256: ec7381e6f759504f54f51a2a6abe9cc42a02699b24eaf722e150e633920c5647

3. Inference engine

Production
awllm_native_minimal · health PASS_SCOPED · readback 30/09/2026
Shadow candidate
AWLLM V2 · PASS_SCOPED 1/1; candidate quality PARTIAL 19/20
Engine ownership
6/31 · PARTIAL; this is component ownership, not model-version ownership
Serving framework
awllm_native_minimal · loaded/healthy; LMDeploy imported: false; call count: 0
Latency / throughput / cost
NOT_RUN as a shared, current, comparable benchmark

receipt_id: AEO_AWLLM_ITERATION_V14_20260816T114034Z · sha256: 789f793b5aa240b18facd7274c3ff28b020e25955594249fd39f4923c6e94ad3

Snapshot lịch sử 30/09/2026

WebChat public/loopback HTTP 200; AWLLM model/tokenizer loaded. MCP gateway 2.2.0: 40 tools; gateway p50 148 ms / p95 4,252 ms / p99 20,047 ms. These are gateway observations, not model performance scores.

runtime receipt: mcp-76a3f2aa2c24a763dfd115b3 · release receipt: mcp-cfeacbda299ea69eb18a688b · full hashes in public manifest live_runtime.receipt_refs

CARM (Capability-Adaptive Research Matrix): LIS/IEC → QIC → CARM → MQR-Mesh ↔ STP → Evidence Reducer → VIF/AWLLM → Progressive Answer → Independent Verify/Credit → LIS/IEC. Architecture canon; current end-to-end research and L1–L8 are NOT_REVERIFIED.

4. AW1 system engine

AEO → VERA/VWR → SEVF/AC → VIF-E/VADR → STREAM/RENDER → WME/ELE → Production readback

AEO
Need gate, capability routing and work orchestration.
VERA / VWR
Canonical answer and workstream renderer; produces one coherent user-facing result from parallel work.
SEVF / AC / ACVL
Evidence admission, atomic-claim verification and claim-level gates. A missing or misbound relation must not be promoted merely because entity tokens appear nearby.
VIF-E / VADR
VIF-E binds source evidence to entity/criterion/predicate scope; VADR carries canonical answer document state through bridge, stream and render lifecycle.
STREAM / RENDER
Public delivery targets one primary surface, one canonical terminal and no post-terminal mutation.
WME / ELE
Improves how verified work and decisions are produced; does not learn blindly from generated answers.
Production readback
Candidate, canary, activation, browser proof and rollback remain separate state transitions.

Architecture names describe system roles. They are not model benchmark scores, and each capability must still be evaluated through its own current receipt.

5. Current release and historical VIF evidence

Current WebChat release: aw1_webchat_v4_17_43_carm_namespace_cleanup_r2_20260930, confirmed by current symlink readback on 30/09/2026. VIF26 is historical evidence from 10/09/2026; its receipt is PASS_SCOPED; this does not promote unrelated UNKNOWN/PARTIAL capabilities.

VIF26 scope
Vietnam Domain Learning · VSIC 36/2025/QĐ-TTg + occupations 34/2020/QĐ-TTg
Classification levels
22 / 87 / 259 / 495 / 743
Core held-out
PASS 8/8
Backend parent
PASS 7/7
UI gates
Primitive UI PASS · Render parity PASS · File UX PASS
Legacy adaptive test
STALE_BASELINE_CONTRACT_NOT_CREDITED
Safety
AWLLM restart: false · tokenizer changed: false · weights changed: false

receipt: VIF26_VIETNAM_DOMAIN_LEARNING_CANDIDATE_RECEIPT · sha256: 7a6e3b88d91d0f115ee8054854f86d6ee7cca9a117c6b65a4ffaa00c4606e95d · generated_at: 2026-09-10T13:18:12Z

Canon boundary. VIF hiện nối tìm → đọc → claim binding → kiểm chứng → Fact/Hypothesis/Unknown → suy luận/frontier → AC/ACVL → answer/stream/render → learning/replay/held-out → memory/skill/training/context → resource optimization → gain/credit/rollback → WME/ELE → next bottleneck. Đây là architecture/canon flow; từng primitive vẫn cần receipt riêng để được gọi là verified.

6. Evaluation and limits

Eval Fabric V4: 24/24 PASS, canonical internal scope measured 17/08/2026; not re-run on the current release. This is not a shared independent frontier benchmark.

7. Reproducibility & citation

Canonical data: aw1-public-manifest.json

@misc{aw1_2026,
  title   = {AW1: Receipt-bound AI Model and System},
  author  = {Nguyen Van Dung and AWAI},
  year    = {2026},
  url     = {https://awai.vn/model-card},
  note    = {Model: sieutocviet/aw1; production inference engine: AWLLM native}
}

truth_epoch: truth-epoch-138276486de37257ffdaacf6 · manifest_sha256: 1a2ae75b757b40a072e8615ea9c0facf4f8aabed63c883ce2e30fe52397b1f84