|
| 1 | +# Iteration-Test Assessment Loops 60–69 |
| 2 | + |
| 3 | +## Scope |
| 4 | + |
| 5 | +Completed ten counted `eforge-assess` loops (60 through 69) against |
| 6 | +`scenarios/iteration-test/scenario.yaml`. Each loop used deterministic scenario/config validation, |
| 7 | +the default non-slow regression suite before generation, forced generation into a new immutable |
| 8 | +loop directory, quantitative evaluation, targeted hard probes, four fresh isolated blind experts, |
| 9 | +and a family-level fix selected from the blind findings. Loop 68 and loop 69 exceeded the mandatory |
| 10 | +disagreement/spread threshold and received fresh report-only deliberations. |
| 11 | + |
| 12 | +Blind agents were restricted to each loop's `output/data` plus the expert briefing/persona |
| 13 | +references. They could not read the scenario, ground truth, evaluation, probes, source, parent |
| 14 | +directories, or prior reports. Dataset integrity manifests confirm the reviewed output did not |
| 15 | +change during each panel. |
| 16 | + |
| 17 | +## Results |
| 18 | + |
| 19 | +| Loop | Records | Eval | Threat / Detection / Network / Host | Initial avg | Deliberation | |
| 20 | +|---:|---:|---:|---|---:|---:| |
| 21 | +| 60 | 80,096 | 97.783 | 66 / 66 / 74 / 94 | 75.00 | — | |
| 22 | +| 61 | 82,379 | 97.086 | 68 / 89 / 82 / 95 | 83.50 | — | |
| 23 | +| 62 | 82,690 | 97.241 | 77 / 86 / 85 / 96 | 86.00 | — | |
| 24 | +| 63 | 82,690 | 97.241 | 77 / 86 / 85 / 96 | 86.00 | — | |
| 25 | +| 64 | 82,676 | 97.241 | 77 / 86 / 73 / 96 | 83.00 | — | |
| 26 | +| 65 | 82,817 | 97.261 | 90 / 86 / 73 / 95 | 86.00 | — | |
| 27 | +| 66 | 80,763 | 97.581 | 74 / 82 / 74 / 91 | 80.25 | — | |
| 28 | +| 67 | 81,573 | 97.403 | 85 / 93 / 71 / 98 | 86.75 | — | |
| 29 | +| 68 | 78,872 | 97.006 | 92 / 70 / 24 / 99 | 71.25 | 91.50 | |
| 30 | +| 69 | 77,567 | 97.291 | 86 / 88 / 28 / 74 | 69.00 | 82.50 | |
| 31 | + |
| 32 | +Loop 69's initial five-loop rolling mean was 78.65. The falling initial mean in loops 68–69 did |
| 33 | +not indicate broad regression: network specialists increasingly judged that subsystem Real, while |
| 34 | +host specialists found sharper endpoint identity contradictions. Deliberation explicitly resolved |
| 35 | +this as scope weighting. |
| 36 | + |
| 37 | +## Family-Level Changes |
| 38 | + |
| 39 | +- **Loop 60:** repaired Windows background-process lifecycle ownership, including durable |
| 40 | + `taskhostw.exe` identity and dependent-aware termination. |
| 41 | +- **Loop 61:** made scenario IP ownership the canonical source for internal PTR identity and private |
| 42 | + reverse-zone authority. |
| 43 | +- **Loop 62:** made retained historical process lifetimes participate in host-local PID reservation. |
| 44 | +- **Loop 63:** added data-driven TLS SNI predicates so domain-specific IDS rules attach only to an |
| 45 | + exact eligible flow. |
| 46 | +- **Loop 64:** made the Linux sudo bundle own the visible `/usr/bin/sudo` lifecycle and PAM PID. |
| 47 | +- **Loop 65:** attached sudo to a live user shell/session, modeled its elevated child, and terminated |
| 48 | + child then sudo after PAM close. |
| 49 | +- **Loop 66:** allowed out-of-order baseline planning to deterministically bootstrap/reuse the sudo |
| 50 | + owner session. |
| 51 | +- **Loop 67:** leased one stable TTY/session shell per host/user/TTY, serialized foreground sudo |
| 52 | + commands, made resolver recovery stateful, and constrained image loads to process lifetime. |
| 53 | +- **Loop 68:** reused historical Linux sessions, unified PAM/eCAR login PID ownership, leased TTYs |
| 54 | + exclusively, restricted `wsqmcons.exe` to rare workstation-only activity, and improved multipart |
| 55 | + curl ownership metadata. |
| 56 | +- **Loop 69 post-review:** resolved Security authentication/Kerberos provider PIDs through the |
| 57 | + host's canonical LSASS identity; registered pre-window sudo sessions as carried state instead of |
| 58 | + visible boundary logins; reused the PAM-owned local-login process for shell ancestry; and made |
| 59 | + child creates a termination floor for the eCAR parent process. |
| 60 | + |
| 61 | +## Final Loop Findings |
| 62 | + |
| 63 | +The loop-69 hard probes verified zero duplicate eCAR event IDs, exclusive TTY ownership across 67 |
| 64 | +successful sudo chains, no module-after-termination rows, and workstation-only `wsqmcons.exe`. |
| 65 | +Twenty-seven of 31 successful PAM local-login records matched the exact eCAR `/bin/login` PID, |
| 66 | +substantially improving loop 68. The blind panel then isolated the remaining two owner seams: |
| 67 | +duplicate login processes/boundary session initialization and literal PID 600 across selected |
| 68 | +Windows Security authentication families. It also confirmed one 28 ms eCAR |
| 69 | +parent-termination-before-child inversion. |
| 70 | + |
| 71 | +Network evidence is the strongest subsystem. The loop-69 network reviewer found coherent |
| 72 | +sensor-local identities for 1,866 shared flows, fully contained protocol/file lifecycles, exact |
| 73 | +loss-aware proxy byte reconciliation, realistic DHCP T/2 renewals, and complete ASA connection/NAT |
| 74 | +lifecycle behavior. The initial network synthetic-confidence score was 28. |
| 75 | + |
| 76 | +## Verification |
| 77 | + |
| 78 | +- Scenario validation: valid with only expected advisory warnings. |
| 79 | +- Config validation: 89 configuration files valid, zero errors. |
| 80 | +- Every generated loop: automated evaluation PASS. |
| 81 | +- Loop-69 quantitative evaluation: 97.29138173504755 over 77,567 records. |
| 82 | +- Focused post-review regression tests: passed. |
| 83 | +- Ruff lint and format checks: passed. |
| 84 | +- Full post-review regression suite: 5,516 passed, 20 skipped in 359.16 seconds; no generated loop |
| 85 | + was mutated after its blind review began. |
| 86 | + |
| 87 | +## Remaining Work |
| 88 | + |
| 89 | +1. Regenerate to verify the post-loop-69 provider-PID and carried-session fixes against output-level |
| 90 | + probes and a new blind panel. |
| 91 | +2. Add coherent source-process visibility/drop semantics for RDP and PsExec initiating endpoints. |
| 92 | +3. Expand public scanner actor and TCP-fingerprint populations with a long one-off tail. |
| 93 | +4. Correct the isolated explicit-credential source-context leak and addressless KDC 4771 request. |
| 94 | +5. Continue reducing local-console density on Linux server roles. |
| 95 | + |
| 96 | +Artifacts are under `scenarios/iteration-test/blind-test/loop-60` through `loop-69`; the aggregate |
| 97 | +report is `scenarios/iteration-test/blind-test/REPORT.md`, and the trend visualization is |
| 98 | +`scenarios/iteration-test/blind-test/assessment-effectiveness-dashboard-last-20-loops.svg`. |
0 commit comments