AW1, tách đúng
model và engine.
AW1 là một hệ thống production gồm model weights, inference engine và agent/evidence engine. Ba lớp có danh tính và mức bằng chứng độc lập; candidate không được trình bày như production.
English summary. AW1 combines the sieutocviet/aw1 model identity, the AWLLM inference runtime, and a receipt-bound research, rendering and verification system. Unpublished fields remain UNKNOWN; this page is not a claim of frontier-model parity.
V2 candidate evidence is historical (16/08/2026); current runtime readback: awllm_native_minimal.
AEO routes work; VERA/VWR owns canonical workstreams; SEVF with AC/ACVL gates evidence; VIF-E/VADR binds evidence to answer lifecycle; WME/ELE credits verified outcomes.
1. Candidate AW — cập nhật 05/10/2026
QUALITY WITHHELD · không promote. Candidate aw1-answerability-entity-ab-20261005-r1 kiểm tra scorer answerability trên QA tổng hợp song ngữ; reader và raw ranking giữ cố định. Đây là thí nghiệm candidate, không phải đánh giá chất lượng production.
Train fitting
Tại checkpoint cold được chọn: control C đạt 8.000/8.000 mỗi ngôn ngữ; reassignment R đạt 9.600/9.600 mỗi ngôn ngữ. Chỉ số này đo quyết định answerability trên dữ liệu fitting, không phải độ chính xác câu trả lời hay held-out.
Development — evidence-group transfer
Hai nhánh C/R có cùng số đếm tại operating point được chọn: EN raw exact 27/40, gated exact 13/40, no-answer đúng 10/20; VI raw exact và gated exact 36/40, no-answer đúng 20/20. Question pairs 13/32; position pairs 19/32. Gate dev FAIL_OR_NOT_RUN; final NOT_RUN_DEV_FAIL; canary NOT_RUN. Không ghi nhận credit cải thiện do reassignment.
Accounting và production
600 cập nhật scorer mỗi nhánh, tổng 1.200 cập nhật answerability; LM updates = 0, reader/start/end updates = 0. Mỗi nhánh có 19.200 binary-label presentations, 16.000 logical rows và 2.000 evidence groups. Candidate chưa chuyển vào production; receipt này không chứng minh chất lượng production.
Điểm chặn: train fitting đã hoàn hảo nhưng không chuyển thành cải thiện answerability trên evidence groups EN mới. QUALITY WITHHELD; không promotion.
Receipt entity A/B 05/10/2026 · SHA256: dff08f355829f4c5496a3875461e74b3d307457cf07967173fdf9497ad795008
Receipt residual A/B trước đó (05/10/2026) có VI gated 35/40, EN gated 13/40, EN no-answer 10/20, question pairs 12/32. Hai receipt dùng packet DEV khác nhau; không coi 35→36/40 là cải thiện đối chứng.
Receipt residual · SHA256: 89d0b2c97bdc8efa79d0b984fc401e4b4cd62b1941ef1b55f616842de86d93ca
Chỉ số lịch sử được công bố ngày 30/09/2026: corpus 1.000.565.587 ký tự từ 6 nhóm nguồn, candidate khoảng 9,3 triệu tham số, 200 bước pilot. Đây là snapshot lịch sử, không cộng 1.200 scorer updates vào LM updates.
2. Historical context contract
65,536 native tokens and 64,884 held-out prompt tokens were published on 03/09/2026; no new long-context test was run for this update.
primary receipt: RuntimeContextFrontierVNext_20260816T224917Z · sha256: ec7381e6f759504f54f51a2a6abe9cc42a02699b24eaf722e150e633920c5647
3. Inference engine
receipt_id: AEO_AWLLM_ITERATION_V14_20260816T114034Z · sha256: 789f793b5aa240b18facd7274c3ff28b020e25955594249fd39f4923c6e94ad3
Snapshot lịch sử 30/09/2026
WebChat public/loopback HTTP 200; AWLLM model/tokenizer loaded. MCP gateway 2.2.0: 40 tools; gateway p50 148 ms / p95 4,252 ms / p99 20,047 ms. These are gateway observations, not model performance scores.
runtime receipt: mcp-76a3f2aa2c24a763dfd115b3 · release receipt: mcp-cfeacbda299ea69eb18a688b · full hashes in public manifest live_runtime.receipt_refs
CARM (Capability-Adaptive Research Matrix): LIS/IEC → QIC → CARM → MQR-Mesh ↔ STP → Evidence Reducer → VIF/AWLLM → Progressive Answer → Independent Verify/Credit → LIS/IEC. Architecture canon; current end-to-end research and L1–L8 are NOT_REVERIFIED.
4. AW1 system engine
AEO → VERA/VWR → SEVF/AC → VIF-E/VADR → STREAM/RENDER → WME/ELE → Production readback
Architecture names describe system roles. They are not model benchmark scores, and each capability must still be evaluated through its own current receipt.
5. Current release and historical VIF evidence
Current WebChat release: aw1_webchat_v4_17_43_carm_namespace_cleanup_r2_20260930, confirmed by current symlink readback on 30/09/2026. VIF26 is historical evidence from 10/09/2026; its receipt is PASS_SCOPED; this does not promote unrelated UNKNOWN/PARTIAL capabilities.
receipt: VIF26_VIETNAM_DOMAIN_LEARNING_CANDIDATE_RECEIPT · sha256: 7a6e3b88d91d0f115ee8054854f86d6ee7cca9a117c6b65a4ffaa00c4606e95d · generated_at: 2026-09-10T13:18:12Z
Canon boundary. VIF hiện nối tìm → đọc → claim binding → kiểm chứng → Fact/Hypothesis/Unknown → suy luận/frontier → AC/ACVL → answer/stream/render → learning/replay/held-out → memory/skill/training/context → resource optimization → gain/credit/rollback → WME/ELE → next bottleneck. Đây là architecture/canon flow; từng primitive vẫn cần receipt riêng để được gọi là verified.
6. Evaluation and limits
Eval Fabric V4: 24/24 PASS, canonical internal scope measured 17/08/2026; not re-run on the current release. This is not a shared independent frontier benchmark.
- No permission to claim AW1 equals or exceeds GPT, Claude, Gemini or Grok.
- Candidate progress is reported separately from production. Final held-out contamination checks, quality evaluation and production promotion remain pending; no independent-production claim is made.
- AWLLM V2 is not production merely because its candidate receipt passed.
- Public manifest snapshot can become stale; always inspect truth epoch, freshness, scope and denominator.
7. Reproducibility & citation
Canonical data: aw1-public-manifest.json
@misc{aw1_2026,
title = {AW1: Receipt-bound AI Model and System},
author = {Nguyen Van Dung and AWAI},
year = {2026},
url = {https://awai.vn/model-card},
note = {Model: sieutocviet/aw1; production inference engine: AWLLM native}
}truth_epoch: truth-epoch-138276486de37257ffdaacf6 · manifest_sha256: 1a2ae75b757b40a072e8615ea9c0facf4f8aabed63c883ce2e30fe52397b1f84