Research: DeepSeek core instrument factorial
Baseline#
Предыдущее исследование
2026-07-26-strong-llm-core-escalation
завершило eight-cell screen:
- DeepSeek deterministic
4/4; - DeepSeek semantic
3/4, total4, один critical grounding failure; - selector:
no_eligible_candidate; - paid finalist calls:
0.
RU scenario сам создал абсолютные immutable facts про constant 1G и запрет
G<1, затем success route разрешил fluctuations/exception и всё равно объявил
compliance. EN scenario той же модели получил semantic 2, поэтому capability
существует, а failure локализован в cross-field semantic consistency.
Метод#
Независимые переменные#
| Factor | Control | Treatment |
|---|---|---|
| Prompt | frozen baseline-v1 |
semantic-v2: один общий silent invariant block |
| Reasoning | off |
medium |
Factorial design измеряет:
- prompt main effect:
semantic-v2противbaseline-v1внутри одинакового reasoning; - reasoning main effect:
mediumпротивoffвнутри одинакового prompt; - interaction: отличается ли combined effect от двух изолированных effects.
Матрица#
- model:
deepseek/deepseek-v4-pro; - role: scenario;
- languages:
ru,en; - seeds:
20260727,20260728,20260729; - configs: четыре preregistered combinations;
- expected paid screening calls:
24.
Historical response не включается в causal comparison: он сохраняется как baseline evidence, но новый control запускается одновременно с treatments, чтобы не спутать intervention с provider drift.
Treatment delta#
semantic-v2:
- требует совместной истинности всех story facts;
- представляет несовместимые testimony/complaints/regulations как fact о наличии conflicting claims, а не как одновременно истинные canonical facts;
- запрещает success через exception/waiver, противоречащий fact;
- требует silent comparison всех facts против summaries, exits, consequences и endings;
- не добавляет output fields и не содержит fixture-specific
1Gподсказок.
Неизменные проверки#
- exact single JSON object;
- AdventureDraft shape;
- authoring validator и compiler;
- bounded structure, IDs, refs, language и content boundary;
- semantic score
0/1/2; - critical grounding flag;
- metadata-hidden isolated review;
- selector hard gates.
Falsifiable expectations#
- Хотя бы один treatment cell отличается от paired control по semantic score или critical flag; иначе intervention effect indistinguishable.
- Eligible config обязан пройти все шесть cells; среднее улучшение не может скрыть единичный hard failure.
- Любой output-format, language, boundary или deterministic regression отклоняет config независимо от semantic average.
Selection и stopping rule#
Config eligible только при:
- six observed raw responses;
- deterministic
6/6; - semantic score ≥1
6/6; - critical grounding failures
0.
Eligible configs ранжируются:
- больше semantic-2 cells;
- выше semantic score total;
- lower intervention order:
baseline-v1--off,semantic-v2--off,baseline-v1--medium,semantic-v2--medium.
Третий пункт выбирает меньшую runtime/maintenance intervention только после полного quality tie. При отсутствии eligible config confirmation запрещена.
Confirmation#
Selected scenario config проверяется на seeds 20260730–20260732.
Intent/narration используют frozen baseline prompt и reasoning off.
Матрица: 2 languages × 3 roles × 3 seeds = 18 calls.
Confirmation pass требует deterministic 18/18, semantic floor во всех
cells и 0 critical failures. First-pass failures не скрываются repair.
Screening results#
Provider batch завершён без потерь: 24/24 jobs ready, 0 pending,
0 failed. Сохранены ровно 24 capture-authoritative raw-файла общим размером
191730 bytes. Frozen source recomputation и blind review приняли полный
набор без уменьшения denominator.
| Config | Observed | Deterministic | Semantic ≥1 | Semantic 2 | Semantic total | Critical |
|---|---|---|---|---|---|---|
baseline-v1--off |
6 | 4/6 | 1/6 | 0 | 1 | 5 |
semantic-v2--off |
6 | 4/6 | 4/6 | 1 | 5 | 2 |
baseline-v1--medium |
6 | 3/6 | 3/6 | 1 | 4 | 3 |
semantic-v2--medium |
6 | 5/6 | 4/6 | 1 | 5 | 2 |
| Effect, treatment minus control | Deterministic pass rate | Semantic-2 rate | Semantic mean | Critical-failure rate |
|---|---|---|---|---|
Prompt main effect, semantic-v2 - baseline-v1 |
+16.7 pp | +8.3 pp | +0.417 | −33.3 pp |
Reasoning main effect, medium - off |
0.0 pp | +8.3 pp | +0.250 | −16.7 pp |
| Interaction | +33.3 pp | −16.7 pp | −0.500 | +33.3 pp |
semantic-v2 улучшил средние метрики, но не устранил единичные hard failures.
Reasoning medium не дал deterministic main effect; положительный
deterministic interaction сопровождался отрицательным semantic interaction.
Frozen selector вернул no_eligible_config. В соответствии со stopping rule
confirmation не запускалась: 0/18, без расхода на запрещённую ветку.
Provider не экспонировал credits, token counts, per-job latency или retry
count. Эти поля записаны как null, без оценок. Возвращённый terminal batch
token отличался от submitted token одним лишним дефисом в неавторитетном
label последнего job; все 24 job IDs, terminal records и output records
совпали exactly. Расхождение сохранено, а не исправлено постфактум.
Evidence registry#
| ID | Тип | Evidence | Использование |
|---|---|---|---|
| E-001 | Measurement | ../2026-07-26-strong-llm-core-escalation/artifacts/derived/finalist-selection.json |
prior terminal decision |
| E-002 | Measurement | prior DeepSeek RU raw | exact contradiction |
| E-003 | Measurement | prior semantic review | unchanged gate baseline |
| E-004 | Fact | Comfy OpenRouter LLM node, accessed 2026-07-26 | DeepSeek V4 Pro supports off/low/medium/high reasoning effort |
| E-005 | Fact | artifacts/derived/protocol-v2.json и identity |
frozen v2 matrix, hashes, unchanged gates и zero-paid supersession v1 |
| E-006 | Measurement | npm test; npm run typecheck |
capture/transport/source/finalization contracts и strict TypeScript |
| E-007 | Measurement | artifacts/derived/screening-preflight-v2-audit.json |
canonical 24-request preflight: 72/72 bindings, 50/50 reproduced files, 0 byte mismatches |
| E-008 | Measurement | artifacts/derived/screening-batch-submission.json |
единственная successful submission, 24 request-to-job bindings |
| E-009 | Measurement | artifacts/derived/screening-batch-capture.json и screening-batch-transport-audit.json |
terminal metadata, exact raw bytes, unavailable telemetry и provider label drift |
| E-010 | Measurement | artifacts/derived/screening-transport/transport.json |
prepared request/workflow/job/raw binding |
| E-011 | Measurement | artifacts/derived/screening-scorecard.json |
frozen deterministic scoring для 24/24 cells |
| E-012 | Measurement | artifacts/derived/screening-semantic-review.json |
независимые metadata-hidden scores и critical flags |
| E-013 | Derived | screening-evidence.json и factorial-analysis.json |
verified join и preregistered main effects/interaction |
| E-014 | Decision | selection.json |
no_eligible_config, confirmation stopped |
Terminal artifact SHA-256:
| Artifact | SHA-256 |
|---|---|
| batch submission | 6b80098c8272e70d9565b58190db6155b50dbdb48d373bb1e6fb6d701fa5c357 |
| batch capture | 253197e4f7e12c639f218e2dbd05d0e0292145a255d804ccde4b31ee661cc832 |
| provider transport audit | 00ee107c4275cd8bbc611adf8251cc2060e71009cc2bc010eba429ef39ed858d |
| transport | f65f9452a33f4c50c27e1c6afa877f7cb94fd0fc0564c68a832a726d11d83b3b |
| scorecard | dac5eb5d71bcd750c4cbabdd3833d0533261be4fdf8a83e9cb7d4f69e492f86c |
| blind semantic review | ed5df980bda88c12af6dffa53e622240dd1e6afde1bdbd39829eec374c72e536 |
| screening evidence | 0d57129271a97b59e0eac9ed96a4922e45b01d380b0694ce1e0d75fbb9bdb074 |
| factorial analysis | 87e5e93c8e7db734a4b5a2b69e02741e8ce7af261226523a28510ef5686f6d22 |
| selection | 1654c760effda3f3cb6f8bb5605f7d241a68b508d54ad8921d261ac5e61b4e73 |
Experiment log#
| Time | Action | Observation | Evidence | Decision |
|---|---|---|---|---|
| 2026-07-26, design | Audited prior RU failure and EN pass | Missing joint-truth/route-consistency instruction is a concrete prompt gap | E-002, E-003 | Test prompt delta without changing scorer |
| 2026-07-26, design | Checked prior workflow and current official node docs | Prior run used reasoning off; DeepSeek Pro supports medium |
E-004 | Use 2×2 factorial instead of assuming prompt-only cause |
| 2026-07-26, user approval | User requested both prompt and reasoning effects | Full factorial approved | conversation | Freeze 24 + conditional 18 matrix |
| 2026-07-26, harness | Capture-authoritative raw, transport-v2, fresh source recomputation, atomic publication and full-chain finalization tested | 139 technical + 5 research-site tests pass; typecheck exit 0 | E-006 | Freeze reviewed implementation |
| 2026-07-26, preflight | Rebuilt screening preparation into a fresh temporary root and compared every byte | 50 expected/50 actual files, 0 mismatches; 72/72 prompt/workflow/payload hashes | E-007 | Freeze protocol v2 against screening-preflight-v2 |
| 2026-07-26, protocol v2 | Pre-paid review closed the evidence-chain gaps before any generation | v2 identity binds protocol, catalog, prepared tree, scorer and rubric; paid calls under v1: 0 | E-005, E-007 | Dry-run only the canonical v2 workflows |
| 2026-07-26, dry-run | Reauthenticated production MCP after an OAuth refresh race, then validated workflows strictly sequentially | 24/24 validated; 0 submitted, 0 jobs, 0 credits | screening-dry-run.json |
Permit one confirmed 24-item paid batch |
| 2026-07-26, paid submit | Submitted the approved batch once with confirm=true |
Definite pre-queue rejection for all 24: Subscription required to queue workflows; no batch ID, no jobs, no credits spent. Authenticated UI showed 1,788 credits but Subscribe to Run |
screening-submission-attempt-20260726T224733+0700.json |
Do not retry; a new real-money subscription needs explicit authority |
| 2026-07-26, subscription | User activated the subscription and explicitly resumed the approved study | Production OAuth authenticated and queue authority restored | conversation, E-008 | Continue only the frozen v2 batch |
| 2026-07-26, paid screening | Queued the canonical workflows in one confirmed batch | 24 unique jobs created; no second submission during OAuth refresh recovery | E-008 | Wait on the owned batch ID |
| 2026-07-26, terminal capture | Waited through the same batch and downloaded signed outputs by job ID | 24 ready, 0 pending, 0 failed; 24 exact files, 191730 bytes | E-009, E-010 | Permit frozen scoring |
| 2026-07-26, deterministic scoring | Recomputed all responses from capture-authoritative bytes | Config pass counts: 4/6, 4/6, 3/6, 5/6 | E-011 | Continue blind review; no config can yet satisfy 6/6 |
| 2026-07-26, blind review | Independent reviewer received only randomized aliases, hashes and raw responses | 24 reviews: score 0 = 12, score 1 = 9, score 2 = 3; 12 critical flags | E-012 | Join only through private source-bound map |
| 2026-07-26, finalization | Reverified the full source chain and applied frozen factorial analysis/selector | no_eligible_config; best deterministic 5/6, best semantic floor 4/6 |
E-013, E-014 | Stop confirmation without weakening gates |
Unknown#
- Actual credit delta, token counts, retries и per-job latency: provider run
завершён, но MCP эти поля не экспонировал; capture хранит
null. - Provider determinism при одинаковом seed hint.
- Human naturalness.
- Genre-diverse и long-session reliability.
- Достаточен ли prompt/reasoning tuning без typed FactLedger.