RnD · Линия доказательств Technical R&D
Каталог и оглавление
TECH.DOC RnD/technical/2026-07-26-deepseek-core-instrument-factorial/RESEARCH.md raw.md ->

Research: DeepSeek core instrument factorial

Baseline#

Предыдущее исследование 2026-07-26-strong-llm-core-escalation завершило eight-cell screen:

  • DeepSeek deterministic 4/4;
  • DeepSeek semantic 3/4, total 4, один critical grounding failure;
  • selector: no_eligible_candidate;
  • paid finalist calls: 0.

RU scenario сам создал абсолютные immutable facts про constant 1G и запрет G<1, затем success route разрешил fluctuations/exception и всё равно объявил compliance. EN scenario той же модели получил semantic 2, поэтому capability существует, а failure локализован в cross-field semantic consistency.

Метод#

Независимые переменные#

Factor Control Treatment
Prompt frozen baseline-v1 semantic-v2: один общий silent invariant block
Reasoning off medium

Factorial design измеряет:

  • prompt main effect: semantic-v2 против baseline-v1 внутри одинакового reasoning;
  • reasoning main effect: medium против off внутри одинакового prompt;
  • interaction: отличается ли combined effect от двух изолированных effects.

Матрица#

  • model: deepseek/deepseek-v4-pro;
  • role: scenario;
  • languages: ru, en;
  • seeds: 20260727, 20260728, 20260729;
  • configs: четыре preregistered combinations;
  • expected paid screening calls: 24.

Historical response не включается в causal comparison: он сохраняется как baseline evidence, но новый control запускается одновременно с treatments, чтобы не спутать intervention с provider drift.

Treatment delta#

semantic-v2:

  1. требует совместной истинности всех story facts;
  2. представляет несовместимые testimony/complaints/regulations как fact о наличии conflicting claims, а не как одновременно истинные canonical facts;
  3. запрещает success через exception/waiver, противоречащий fact;
  4. требует silent comparison всех facts против summaries, exits, consequences и endings;
  5. не добавляет output fields и не содержит fixture-specific 1G подсказок.

Неизменные проверки#

  • exact single JSON object;
  • AdventureDraft shape;
  • authoring validator и compiler;
  • bounded structure, IDs, refs, language и content boundary;
  • semantic score 0/1/2;
  • critical grounding flag;
  • metadata-hidden isolated review;
  • selector hard gates.

Falsifiable expectations#

  1. Хотя бы один treatment cell отличается от paired control по semantic score или critical flag; иначе intervention effect indistinguishable.
  2. Eligible config обязан пройти все шесть cells; среднее улучшение не может скрыть единичный hard failure.
  3. Любой output-format, language, boundary или deterministic regression отклоняет config независимо от semantic average.

Selection и stopping rule#

Config eligible только при:

  • six observed raw responses;
  • deterministic 6/6;
  • semantic score ≥1 6/6;
  • critical grounding failures 0.

Eligible configs ранжируются:

  1. больше semantic-2 cells;
  2. выше semantic score total;
  3. lower intervention order: baseline-v1--off, semantic-v2--off, baseline-v1--medium, semantic-v2--medium.

Третий пункт выбирает меньшую runtime/maintenance intervention только после полного quality tie. При отсутствии eligible config confirmation запрещена.

Confirmation#

Selected scenario config проверяется на seeds 2026073020260732. Intent/narration используют frozen baseline prompt и reasoning off. Матрица: 2 languages × 3 roles × 3 seeds = 18 calls.

Confirmation pass требует deterministic 18/18, semantic floor во всех cells и 0 critical failures. First-pass failures не скрываются repair.

Screening results#

Provider batch завершён без потерь: 24/24 jobs ready, 0 pending, 0 failed. Сохранены ровно 24 capture-authoritative raw-файла общим размером 191730 bytes. Frozen source recomputation и blind review приняли полный набор без уменьшения denominator.

Config Observed Deterministic Semantic ≥1 Semantic 2 Semantic total Critical
baseline-v1--off 6 4/6 1/6 0 1 5
semantic-v2--off 6 4/6 4/6 1 5 2
baseline-v1--medium 6 3/6 3/6 1 4 3
semantic-v2--medium 6 5/6 4/6 1 5 2
Effect, treatment minus control Deterministic pass rate Semantic-2 rate Semantic mean Critical-failure rate
Prompt main effect, semantic-v2 - baseline-v1 +16.7 pp +8.3 pp +0.417 −33.3 pp
Reasoning main effect, medium - off 0.0 pp +8.3 pp +0.250 −16.7 pp
Interaction +33.3 pp −16.7 pp −0.500 +33.3 pp

semantic-v2 улучшил средние метрики, но не устранил единичные hard failures. Reasoning medium не дал deterministic main effect; положительный deterministic interaction сопровождался отрицательным semantic interaction. Frozen selector вернул no_eligible_config. В соответствии со stopping rule confirmation не запускалась: 0/18, без расхода на запрещённую ветку.

Provider не экспонировал credits, token counts, per-job latency или retry count. Эти поля записаны как null, без оценок. Возвращённый terminal batch token отличался от submitted token одним лишним дефисом в неавторитетном label последнего job; все 24 job IDs, terminal records и output records совпали exactly. Расхождение сохранено, а не исправлено постфактум.

Evidence registry#

ID Тип Evidence Использование
E-001 Measurement ../2026-07-26-strong-llm-core-escalation/artifacts/derived/finalist-selection.json prior terminal decision
E-002 Measurement prior DeepSeek RU raw exact contradiction
E-003 Measurement prior semantic review unchanged gate baseline
E-004 Fact Comfy OpenRouter LLM node, accessed 2026-07-26 DeepSeek V4 Pro supports off/low/medium/high reasoning effort
E-005 Fact artifacts/derived/protocol-v2.json и identity frozen v2 matrix, hashes, unchanged gates и zero-paid supersession v1
E-006 Measurement npm test; npm run typecheck capture/transport/source/finalization contracts и strict TypeScript
E-007 Measurement artifacts/derived/screening-preflight-v2-audit.json canonical 24-request preflight: 72/72 bindings, 50/50 reproduced files, 0 byte mismatches
E-008 Measurement artifacts/derived/screening-batch-submission.json единственная successful submission, 24 request-to-job bindings
E-009 Measurement artifacts/derived/screening-batch-capture.json и screening-batch-transport-audit.json terminal metadata, exact raw bytes, unavailable telemetry и provider label drift
E-010 Measurement artifacts/derived/screening-transport/transport.json prepared request/workflow/job/raw binding
E-011 Measurement artifacts/derived/screening-scorecard.json frozen deterministic scoring для 24/24 cells
E-012 Measurement artifacts/derived/screening-semantic-review.json независимые metadata-hidden scores и critical flags
E-013 Derived screening-evidence.json и factorial-analysis.json verified join и preregistered main effects/interaction
E-014 Decision selection.json no_eligible_config, confirmation stopped

Terminal artifact SHA-256:

Artifact SHA-256
batch submission 6b80098c8272e70d9565b58190db6155b50dbdb48d373bb1e6fb6d701fa5c357
batch capture 253197e4f7e12c639f218e2dbd05d0e0292145a255d804ccde4b31ee661cc832
provider transport audit 00ee107c4275cd8bbc611adf8251cc2060e71009cc2bc010eba429ef39ed858d
transport f65f9452a33f4c50c27e1c6afa877f7cb94fd0fc0564c68a832a726d11d83b3b
scorecard dac5eb5d71bcd750c4cbabdd3833d0533261be4fdf8a83e9cb7d4f69e492f86c
blind semantic review ed5df980bda88c12af6dffa53e622240dd1e6afde1bdbd39829eec374c72e536
screening evidence 0d57129271a97b59e0eac9ed96a4922e45b01d380b0694ce1e0d75fbb9bdb074
factorial analysis 87e5e93c8e7db734a4b5a2b69e02741e8ce7af261226523a28510ef5686f6d22
selection 1654c760effda3f3cb6f8bb5605f7d241a68b508d54ad8921d261ac5e61b4e73

Experiment log#

Time Action Observation Evidence Decision
2026-07-26, design Audited prior RU failure and EN pass Missing joint-truth/route-consistency instruction is a concrete prompt gap E-002, E-003 Test prompt delta without changing scorer
2026-07-26, design Checked prior workflow and current official node docs Prior run used reasoning off; DeepSeek Pro supports medium E-004 Use 2×2 factorial instead of assuming prompt-only cause
2026-07-26, user approval User requested both prompt and reasoning effects Full factorial approved conversation Freeze 24 + conditional 18 matrix
2026-07-26, harness Capture-authoritative raw, transport-v2, fresh source recomputation, atomic publication and full-chain finalization tested 139 technical + 5 research-site tests pass; typecheck exit 0 E-006 Freeze reviewed implementation
2026-07-26, preflight Rebuilt screening preparation into a fresh temporary root and compared every byte 50 expected/50 actual files, 0 mismatches; 72/72 prompt/workflow/payload hashes E-007 Freeze protocol v2 against screening-preflight-v2
2026-07-26, protocol v2 Pre-paid review closed the evidence-chain gaps before any generation v2 identity binds protocol, catalog, prepared tree, scorer and rubric; paid calls under v1: 0 E-005, E-007 Dry-run only the canonical v2 workflows
2026-07-26, dry-run Reauthenticated production MCP after an OAuth refresh race, then validated workflows strictly sequentially 24/24 validated; 0 submitted, 0 jobs, 0 credits screening-dry-run.json Permit one confirmed 24-item paid batch
2026-07-26, paid submit Submitted the approved batch once with confirm=true Definite pre-queue rejection for all 24: Subscription required to queue workflows; no batch ID, no jobs, no credits spent. Authenticated UI showed 1,788 credits but Subscribe to Run screening-submission-attempt-20260726T224733+0700.json Do not retry; a new real-money subscription needs explicit authority
2026-07-26, subscription User activated the subscription and explicitly resumed the approved study Production OAuth authenticated and queue authority restored conversation, E-008 Continue only the frozen v2 batch
2026-07-26, paid screening Queued the canonical workflows in one confirmed batch 24 unique jobs created; no second submission during OAuth refresh recovery E-008 Wait on the owned batch ID
2026-07-26, terminal capture Waited through the same batch and downloaded signed outputs by job ID 24 ready, 0 pending, 0 failed; 24 exact files, 191730 bytes E-009, E-010 Permit frozen scoring
2026-07-26, deterministic scoring Recomputed all responses from capture-authoritative bytes Config pass counts: 4/6, 4/6, 3/6, 5/6 E-011 Continue blind review; no config can yet satisfy 6/6
2026-07-26, blind review Independent reviewer received only randomized aliases, hashes and raw responses 24 reviews: score 0 = 12, score 1 = 9, score 2 = 3; 12 critical flags E-012 Join only through private source-bound map
2026-07-26, finalization Reverified the full source chain and applied frozen factorial analysis/selector no_eligible_config; best deterministic 5/6, best semantic floor 4/6 E-013, E-014 Stop confirmation without weakening gates

Unknown#

  • Actual credit delta, token counts, retries и per-job latency: provider run завершён, но MCP эти поля не экспонировал; capture хранит null.
  • Provider determinism при одинаковом seed hint.
  • Human naturalness.
  • Genre-diverse и long-session reliability.
  • Достаточен ли prompt/reasoning tuning без typed FactLedger.