Reference: See neural-gnn/paper/main.pdf for current context understanding.
Map the simulation-GNN training landscape: understand which neural activity simulation configurations allow successful GNN training (connectivity_R2 > 0.9) and which are fundamentally harder. In particular is it possible to use GNN to recover connectivity from low-rank data, from low-gain data, from sparse data, from noisy dynamics ? It is taken for granted that increasing training size (number frames) improves recovery of the neural dynamics parameter. Also increasing the overall complexity of the training data, e.g. its effictive rank, improves recovery. Finally we observe that injecting noise into the simulation (not measurement noise) helps recovery too. This can be however reconsidered. And we need to find other way to make GNN as an efficient tool for neuron recovery.
DO NOT change n_neurons under any circumstances. The n_neurons=1000 value is LOCKED for this exploration. Dedicated LLM-loops are running separately for n=200 low_rank and sparse regimes.
- n_neurons = 1000 is LOCKED - DO NOT modify n_neurons under any circumstances
- FOCUS on n_neurons=1000, with 1 or 4 neuron types, try 2 hours computation
- dedicated LLM-loops have been launched separately to investigate (n=200,low_rank=20) and (n=200, filling=50%)
Each block = n_iter_block iterations exploring one simulation configuration.
The prompt provides: Block info: block {block_number}, iteration {iter_in_block}/{n_iter_block} within block
You maintain TWO files:
File: {config}_analysis.md
- Append every iteration's full log entry
- Append block summaries
- Never read this file — it's for human record only
File: {config}_memory.md
- READ at start of each iteration
- UPDATE at end of each iteration
- Contains: established principles + previous blocks summary + current block iterations
- Fixed size (~500 lines max)
Read {config}_memory.md to recall:
- Established principles
- Previous block findings
- Understanding of simulation-GNN training landscape
- Current block progress
Metrics from analysis.log:
effective rank (99% var): CRITICAL - SVD rank at 99% cumulative variance. Extract this value and log it aseff_rank=Nin the activity field.test_R2: R² between ground truth and rollout predictiontest_pearson: Pearson correlation between ground truth and rollout predictionconnectivity_R2: R² of learned vs true connectivity weightsfinal_loss: final training losscluster_accuracy: neuron classificationkinograph_R2: mean per-frame R² between GT and GNN rollout kinographskinograph_SSIM: structural similarity between kinographskinograph_Wasserstein: time-unaligned population mode Wasserstein (PCA-projected, normalized by GT std — dimensionless; 0 = identical modes, 1 = one-σ shift; captures right modes even if timing differs)
Example analysis.log format:
spectral radius: 1.029
--- activity ---
effective rank (90% var): 26
effective rank (99% var): 84 <-- Extract this value for eff_rank
Classification:
- Converged: connectivity_R2 > 0.9
- Partial: connectivity_R2 0.1-0.9
- Failed: connectivity_R2 < 0.1
Degeneracy Detection (CRITICAL — check every iteration):
Degeneracy occurs when the GNN finds a connectivity matrix W that produces correct activity dynamics but does NOT match the ground-truth W. The MLPs (lin_edge, lin_phi) compensate for wrong W by reshaping their nonlinear mappings. This is a fundamental inverse-problem ambiguity: multiple W can generate indistinguishable time series.
Compute the degeneracy gap = test_pearson - connectivity_R2:
| test_pearson | connectivity_R2 | Degeneracy gap | Diagnosis |
|---|---|---|---|
| > 0.95 | > 0.9 | < 0.1 | Healthy — good dynamics from correct W |
| > 0.95 | 0.3–0.9 | 0.1–0.7 | Degenerate — good dynamics from wrong W |
| > 0.95 | < 0.3 | > 0.7 | Severely degenerate — MLPs fully compensating |
| < 0.5 | < 0.5 | ~0 | Failed — both dynamics and W poor (not degeneracy) |
When degeneracy gap > 0.3, DO NOT trust dynamics metrics as evidence of learning quality. The model has learned to mimic traces, not to recover connectivity. Parameter tuning (lr_W, L1, batch_size) alone cannot resolve degeneracy — it is a structural problem.
Regimes prone to degeneracy:
- Sparse connectivity (filling_factor < 1): many zero weights give the optimizer freedom to redistribute non-zero entries
- Subcritical spectral radius (rho < 1): decaying dynamics carry less information about W structure
- Low eff_rank: fewer distinguishable modes → more equivalent W solutions
How to detect degeneracy in practice:
- Check
test_pearsonvsconnectivity_R2every iteration — if pearson is high but conn is low, report degeneracy - Check if connectivity_R2 plateaus while test_pearson stays at ~1.0 — the GNN has converged to a degenerate basin
- In the activity plots: if rollout traces look qualitatively correct but the W heatmap differs from ground truth → degeneracy
Log degeneracy in the iteration entry when detected:
Degeneracy: gap=0.53 (test_pearson=0.999, conn_R2=0.466) — MLP compensation suspected
Known interventions that can help break degeneracy:
- Increase coeff_edge_diff (100 → 500): monotonicity regularizer constrains lin_edge, reducing its ability to compensate for wrong W
- Increase n_frames (10k → 30k–50k): more diverse activity patterns constrain the solution space
- Increase L1 on W (1E-5 → 1E-4): if true W is sparse, stronger L1 pushes toward sparse solutions that match the true structure
- Increase n_epochs (2 → 3+): more training passes help escape degenerate basins
- The LLM should propose and test additional strategies based on the specific regime
Key insight from Block 7: At sparse 50%, conn_R2 plateaued at 0.466 despite test_pearson=0.999 across 12 iterations. Parameter sweeps (lr_W: 2E-3→1E-2, n_epochs: 1→3, L1: 1E-5→1E-6) improved conn_R2 by +0.29 but could not close the degeneracy gap. This indicates a structural data limit at 10k frames for sparse subcritical networks.
Visual Analysis of Embedding Plot (ONLY if n_neuron_types > 1):
Examine the latest embedding plot to assess cluster quality:
- Location:
log/signal/{config_name}/tmp_training/embedding/ - Files:
{epoch}_{iteration}.png— find the one with highest iteration number
What to look for:
Well-separated clusters (one color per neuron type):
- Each neuron type should form a distinct, tight cluster
- Colors should not overlap significantly
- This indicates embeddings are learning neuron-type information
Log the embedding observation in the iteration log as:
Embedding: [description, e.g., "4 well-separated clusters" or "colors mixed, no separation" or "2 clusters merged"]
Dual-Objective Optimization (when n_neuron_types > 1):
When connectivity_R2 > 0.9 but cluster_accuracy < 0.9:
- Focus on embedding optimization: increase lr_emb, adjust training duration
- The connectivity learning is successful, now optimize embedding learning
| connectivity_R2 | cluster_accuracy | Status | Action |
|---|---|---|---|
| > 0.9 | > 0.9 | FULL CONVERGED | Success - both objectives met |
| > 0.9 | < 0.9 | W-converged | Focus on lr_emb, embedding learning |
| < 0.9 | > 0.9 | E-converged | Focus on lr_W, connectivity learning |
| < 0.9 | < 0.9 | Partial/Failed | Standard optimization |
Upper Confidence Bound (UCB) scores from ucb_scores.txt:
- Provides computed UCB scores for all exploration nodes within a block including current iteration
- At block boundaries, the UCB file will be empty (erased). When empty, use
parent=root
Example:
Node 2: UCB=2.175, parent=1, visits=1, R2=0.997
Node 1: UCB=2.110, parent=root, visits=2, R2=0.934
Append to Full Log ({config}_analysis.md) and Current Block sections of {config}_memory.md:
- In memory.md: Insert iteration log in "Iterations This Block" section (BEFORE "Emerging Observations")
- Update "Emerging Observations" at the END of the file
Log Form:
## Iter N: [converged/partial/failed]
Node: id=N, parent=P
Mode/Strategy: [success-exploit/failure-probe]/[exploit/explore/boundary]
Config: lr_W=X, lr=Y, lr_emb=Z, coeff_W_L1=W, batch_size=B, low_rank_factorization=[T/F], low_rank=R, n_frames=NF, recurrent=[T/F], time_step=T
Metrics: test_R2=A, test_pearson=B, connectivity_R2=C, cluster_accuracy=D, final_loss=E, kino_R2=F, kino_SSIM=G, kino_WD=H
Activity: eff_rank=R (from analysis.log "effective rank (99% var)"), spectral_radius=S, [brief description]
Embedding: [visual observation from embedding plot when n_types>1, e.g., "4 well-separated clusters" or "colors mixed"]
Mutation: [param]: [old] -> [new]
Parent rule: [one line]
Observation: [one line]
Next: parent=P
CRITICAL: Always extract effective rank (99% var) from analysis.log and include it in the Activity field as eff_rank=R. This value is essential for understanding training difficulty and must be recorded in every iteration.
CRITICAL: The Mutation: line is parsed by the UCB tree builder. Always include the exact parameter change with old and new values (e.g., Mutation: lr_W: 2E-3 -> 5E-3). This line MUST appear in the analysis.md log entry for every iteration. Without it, the UCB tree cannot track what was changed.
CRITICAL: When n_neuron_types > 1, always examine the embedding plot and include the Embedding: field in the log. Visual inspection of cluster separation is essential for understanding embedding learning quality.
Step A: Select parent node
- Read
ucb_scores.txt - If empty →
parent=root - Otherwise → select node with highest UCB as parent
CRITICAL: The parent=P in the Node line must be the node ID (integer) of the selected parent, NOT "root" (unless UCB file is empty). Example: if you select node 3 as parent, write Node: id=4, parent=3.
Step B: Choose strategy
| Condition | Strategy | Action |
|---|---|---|
| Default | exploit | Highest UCB node, try mutation |
| 3+ consecutive R² ≥ 0.9 | failure-probe | Extreme parameter to find boundary |
| n_iter_block/4 consecutive successes | explore | Select outside recent chain |
| 2+ distant nodes with R² > 0.9 | recombine | Merge params from both nodes |
| 4+ consecutive partial (0.1 < R² < 0.9) with improving trend | scale-up | Increase data_augmentation_loop (5x) to break plateau |
| 4+ consecutive converged (R²>0.9) with same param dimension | forced-branch | Select 2nd highest UCB node (not recent chain), switch param dimension |
| degeneracy gap > 0.3 for 3+ consecutive iters AND conn_R2 plateaued | degeneracy-break | MLP compensating for wrong W — try: (1) coeff_edge_diff 100→500, (2) recurrent time_step=4-16, (3) increase n_frames, (4) increase L1. Do NOT keep tuning lr_W/lr — parameter sweeps cannot resolve structural degeneracy |
| eff_rank < 8 AND spectral_radius < 0.7 AND 3+ consecutive R² < 0.05 | fixed-point-collapse | Data fundamentally unrecoverable - skip remaining iterations, modify sim params at block end |
| n_neurons ≥ 500 AND 4+ consecutive partial AND n_epochs < 3 | epoch-scale-up | n≥500 needs more training capacity - increase n_epochs to 3+ |
| new regime block AND block 1 convergence rate > 80% | regime-transfer-test | first batch of new regime: test if block 1 optimal params transfer; vary the regime-specific param (e.g. factorization for low-rank) |
| low_rank regime AND no factorization tested | factorization-probe | always test low_rank_factorization=True vs False early in low-rank blocks |
| lr_W increased AND dynamics degraded (test_R2 dropped) | lr-co-optimize | when lr_W increase hurts test_R2, try increasing lr proportionally (lr_W/lr ratio ~10:1) |
| low_rank_factorization=True AND lr_W < 5E-3 | factorization-lr-guard | factorization requires lr_W>=5E-3 to optimize W=W_L@W_R; do NOT use factorization with lr_W<5E-3 |
| lr increased AND dynamics degraded at moderate lr_W (<=3E-3) | lr-ceiling | at moderate lr_W (<=3E-3), lr=1E-4 is optimal; do NOT increase lr beyond 1E-4 at lr_W<=3E-3 |
| lr increased AND dynamics degraded at any lr_W | lr-ceiling-global | lr=1E-4 is optimal in chaotic (eff_rank=35); EXCEPTION 1: in low eff_rank regimes (eff_rank~12, e.g. Dale_law or low_rank), lr=2E-4 may work without degradation (iter 36); EXCEPTION 2: in noisy regimes (eff_rank>=84), lr=2E-4 safe (iters 56, 60); test in context |
| batch_size increased AND dynamics degraded | batch-size-trade-off | batch_size=16 may degrade dynamics ~2% at L1=1E-5 but shows no degradation at L1=1E-6 (iter 24); BUT in Dale regime, batch=16 degrades connectivity ~10% (iter 34); regime-dependent |
| low eff_rank regime AND coeff_W_L1 >= 1E-5 AND dynamics degraded | L1-reduction | in low eff_rank regimes (eff_rank<20), L1=1E-6 is critical for dynamics in low_rank; in Dale regime L1 effect is marginal (iter 35 vs 31); always try L1=1E-6 early in low-rank blocks |
| new regime block AND previous block found L1=1E-6 beneficial | L1-transfer-test | test L1=1E-6 in new regime to see if L1 reduction benefit transfers across regimes |
| Dale_law=True AND lr_W >= 5E-3 | Dale-lr_W-cliff | Dale_law creates sharp lr_W cliff at 4.5E-3-5E-3; lr_W>=5E-3 fails reproducibly (iters 29, 30); safe range [3.5E-3, 4.5E-3]; do NOT use lr_W>=5E-3 with Dale_law |
| eff_rank <= 12 AND constraints (Dale/low_rank) AND lr_W > 4.5E-3 | constrained-lr_W-guard | constrained regimes (Dale, low_rank) narrow the safe lr_W range; always probe cliff before exploiting high lr_W |
| n_neuron_types > 1 AND coeff_W_L1 >= 1E-5 AND n_neurons <= 100 | heterogeneous-L1-guard | L1=1E-6 critical for embedding at n=100/4types (L1=1E-5 destroys cluster_acc 0.990->0.440, iter 44); BUT at n=200/4types L1=1E-5 is BETTER (cluster=1.000 vs 0.750, iter 149 vs 151); n-dependent L1 overrides heterogeneous rule at n>=200 |
| n_neuron_types > 1 AND lr_emb=1E-3 AND lr_W < 4E-3 | lr_emb-lr_W-coupling | lr_emb=1E-3 overshoots at lr_W<4E-3 (cluster 0.710 at lr_W=2E-3, iter 42); use lr_emb=5E-4 when lr_W<=3E-3; lr_emb/lr_W ratio ~0.2 is safe |
| n_neuron_types > 1 AND n_neurons >= 200 AND lr_emb > 1E-3 | heterogeneous-lr_emb-ceiling | lr_emb=2E-3 overshoots at n=200/4types/lr_W=8E-3: cluster 1.000→0.375, dynamics -18.2% (iter 154); lr_emb=1E-3 is ceiling at n>=200; lr_emb/lr_W ratio must be <=0.125 (stricter than n=100 ratio ~0.2) |
| n_neuron_types > 1 AND lr_W > 5E-3 AND n_neurons <= 100 | heterogeneous-lr_W-cap-n100 | at n=100/4types: lr_W=6E-3 degrades embedding (1.000->0.490, iter 45); lr_W=5E-3 sweet spot; do NOT exceed 5E-3 |
| n_neuron_types > 1 AND n_neurons >= 200 AND lr_W > 8E-3 | heterogeneous-lr_W-cap-n200 | at n=200/4types: lr_W=8E-3 is optimal for full dual (conn=0.988, cluster=1.000, iter 149); lr_W=1E-2 DEGRADES (-8.7% conn, -75.5pp cluster, iter 153); lr_W=1.2E-2 partial recovery (conn=0.955, cluster=0.750, iter 155) — NON-MONOTONIC; safe range [6E-3, 8E-3]; do NOT use lr_W=1E-2 |
| n_neuron_types > 1 AND batch_size > 8 | heterogeneous-batch-guard | batch_size=16 severely degrades heterogeneous networks — dynamics -12%, embedding -50%, connectivity -3% (iter 48); always use batch_size=8 for n_types>1 |
| noise_model_level >= 0.5 AND lr_W >= 8E-3 | noisy-lr_W-ceiling | in noisy regimes (noise>=0.5), lr_W=1E-2 degrades dynamics severely (test_R2=0.707, iter 57); optimal lr_W inversely correlates with noise: noise=0.1→2E-3, noise=0.5→4E-3, noise=1.0→2E-3; keep lr_W<=8E-3 at noise>=0.5 |
| noise_model_level > 0 AND lr increased beyond 1E-4 | noisy-lr-tolerance | at noise-inflated eff_rank>=84, lr=2E-4 does NOT degrade dynamics (iter 56, 60); noise widens lr tolerance; OVERRIDES lr-ceiling-global in noisy regimes |
| n_neurons >= 200 AND lr_W < 5E-3 | n-scaling-lr_W-boundary | convergence boundary scales |
| n_neurons >= 200 AND lr_W > optimal_lr_W | n-scaling-dynamics-cliff | dynamics cliff depends on n AND n_epochs: at 1ep n=100→8E-3, n=200→5.5E-3, n=300→~1.2E-2; at 2ep n=200 cliff shifts to >1.2E-2 (all converge up to 1.2E-2, iter 131); more epochs widen safe lr_W range; n=200/2ep optimal at lr_W=8E-3 (iter 126); n=300/3ep optimal at lr_W=1E-2 |
| n_neurons >= 200 AND n_epochs >= 2 AND lr_W in [5E-3, 1.2E-2] | n200-epoch-recipe | n=200 at 2ep achieves 100% convergence (12/12, block 11); optimal: lr_W=8E-3, lr=2E-4, L1=1E-5; 3ep further boosts conn 0.993→0.994 and dynamics 0.963→0.985; batch=16 safe (negligible degradation); do NOT use L1=1E-6 at n=200 |
| n_neurons >= 200 AND lr > 2E-4 AND lr_W >= 1E-2 | lr-lr_W-interaction | lr tolerance narrows at high lr_W: at n=300/lr_W=1E-2, lr=3E-4 degrades dynamics -4.9% and conn -2.9% (iter 108); lr=2E-4 is safe; at n=200/lr_W=5E-3, lr=3E-4 was safe; higher lr_W tightens lr ceiling; use lr=2E-4 when lr_W>=1E-2 |
| n_neurons >= 200 AND lr <= 1E-4 | n-scaling-lr-tolerance | n=200 (eff_rank=43) tolerates lr=2-3E-4 at lr_W=5E-3 (iters 67, 72); n=300 tolerates lr=2E-4 at lr_W=1E-2 but NOT lr=3E-4 (iter 108); lr tolerance depends on lr_W — safe at 2E-4 for all n>=200 |
| n_types=1 AND chaotic AND coeff_W_L1 < 1E-5 AND n_neurons <= 200 | L1-chaotic-homogeneous-guard | L1=1E-6 harmful for n_types=1 chaotic at n<=200: degrades dynamics; at n=300 L1=1E-6 BENEFICIAL (+4.1% conn, iter 119); at n=600 L1=1E-6 HARMFUL AGAIN (-3.4% conn, iter 142); L1=1E-6 benefit is NARROW to n=300 only; use L1=1E-5 for n<=200 and n>=600, L1=1E-6 for n=300 |
| n_neurons >= 300 AND n_epochs < 3 AND coeff_W_L1 >= 1E-5 | n300-epoch-L1-synergy | n=300 convergence requires BOTH n_epochs>=3 AND L1=1E-6: n_epochs=3+L1=1E-5→0.886 (partial); n_epochs=3+L1=1E-6→0.922 (converged); n_epochs=4+L1=1E-6→0.924 (best); L1=1E-6 contributes +4.1% and n_epochs contributes +3%; both needed for convergence (iters 110, 114, 117, 119) |
| n_neurons >= 300 AND n_epochs < 2 | n300-epoch-minimum | n=300 at 10k frames requires n_epochs>=2: best conn at 1 epoch was 0.805 (iter 103), n_epochs=2 boosted to 0.890 (+10.6%, iter 106); loss dropped 4x; n=300 is training-capacity-limited at 1 epoch; use n_epochs>=2 for n>=300 |
| n_neurons >= 300 AND batch_size > 8 | n300-batch-guard | batch_size=16 degrades conn -7.8% at n=300 (0.893→0.823, iter 115); extends heterogeneous-batch-guard to large n; always use batch_size=8 for n>=300 |
| connectivity_filling_factor < 1 AND n_epochs < 2 | sparse-epoch-minimum | sparse connectivity (filling_factor=0.5) at 10k frames requires n_epochs>=2; at 1 epoch no config converges (max R2=0.420); 2 epochs boosts ~10% (0.466); lr_W=1E-2 optimal (no cliff); eff_rank drops to 21, spectral_radius=0.746 subcritical (iters 73-84) |
| connectivity_filling_factor < 1 AND all partial for full block | sparse-scale-up | sparse 50% at n=100/10k frames/2 epochs yields max R2=0.466 (0% convergence); reference config uses n=1000/100k frames/10 epochs — sparse regime fundamentally needs more data and training to converge; consider n_frames increase at block boundaries |
| connectivity_filling_factor < 1 AND eff_rank < 25 | sparse-subcritical-guard | sparse connectivity reduces eff_rank (35->21 at 50% fill) and makes spectral_radius subcritical (0.746); dynamics test_R2 stuck at ~0.11 regardless of training params; this is a data complexity issue not a training issue |
| connectivity_filling_factor < 1 AND noise_model_level > 0 AND conn_R2 plateau for 4+ iters | sparse-noise-plateau | noise inflates eff_rank (21→91 at noise=0.5) but does NOT rescue sparse connectivity; conn plateaus at ~0.489 with COMPLETE parameter insensitivity (lr_W 2E-3 to 1.5E-2, L1, n_epochs, aug_loop all identical); subcritical rho=0.746 unchanged by noise; sparse+noise is data-limited not training-limited (iters 85-96) |
| recurrent_training=True AND spectral_radius < 1 AND noise_model_level > 0 | recurrent-subcritical-guard | recurrent training (time_step=4) is CATASTROPHIC in noisy subcritical regimes: conn collapsed 0.489→0.054, loss exploded 200→170k (iter 95); multi-step rollout amplifies instability when dynamics are subcritical + noisy; do NOT use recurrent training when spectral_radius<1 and noise>0 |
| n_neurons >= 600 AND lr <= 1E-4 AND n_frames <= 10000 | n600-lr-floor | lr=1E-4 is CATASTROPHIC at n=600/10k: conn=0.000 (iter 143); BUT at n=600/30k lr=1E-4 gives conn=0.993 (iter 192) — n_frames rescues lr=1E-4; this rule applies ONLY at n_frames<=10k |
| n_neurons >= 600 AND n_epochs < 10 AND n_frames <= 10000 | n600-epoch-minimum | n=600 at 10k frames needs n_epochs>=10: 4ep→0.540, 10ep→0.626; BUT at 30k frames 2ep→0.967, 5ep→0.993 (block 16); this rule applies ONLY at n_frames<=10k |
| n_neurons >= 600 AND coeff_W_L1 < 1E-5 AND n_frames <= 10000 | n600-L1-guard | L1=1E-6 worse than L1=1E-5 at n=600/10k; BUT at 30k L1 sensitivity vanishes (iter 188: L1=1E-6 marginally better +0.6%); this rule applies ONLY at n_frames<=10k |
| n_neurons >= 600 AND lr_W > 1E-2 AND n_frames <= 10000 | n600-lr_W-ceiling-10k | at 10k: lr_W=1.5E-2 and 2E-2 worse (0.684); lr_W=1E-2 optimal at n=600/10k |
| n_neurons >= 600 AND n_frames >= 30000 | n600-30k-recipe | n_frames=30k transforms n=600: 0%→100% convergence; eff_rank 50→87; optimal lr_W=5E-3 (conn=0.993, dynamics=0.966 at 5ep); lr_W range [3E-3, 1.5E-2] all converge; batch=16 safe; L1 insensitive; 2ep sufficient; dynamics-optimal lr_W shifts to 5E-3 (lower than 10k's 1E-2); lr=1E-4 NOT catastrophic |
| recurrent_training=True AND spectral_radius >= 1 | recurrent-supercritical-tradeoff | recurrent training (time_step=4) at supercritical rho=1.064: conn improves +0.3% (0.990→0.993) but dynamics degrade -12.3% (test_R2 0.934→0.819); NOT Pareto-improving — trade-off between W recovery and dynamics quality; recurrent may help in regimes where conn is the bottleneck (e.g. n>=300) |
| recurrent_training=True AND noise_recurrent_level >= 0.05 | recurrent-noise-ceiling | noise_recurrent_level=0.05 degrades conn from 0.993→0.772 (partial) and loss 28x higher (iter 168); noise_rec=0.01 tolerable with warmup; keep noise_recurrent_level<=0.01 |
| recurrent_training=True AND recurrent_training_start_epoch >= 1 | recurrent-warmup-tradeoff | warmup (start_ep=1) recovers dynamics +9.5% but conn drops -8.2% (iter 167 vs 166); warmup shifts capacity from W toward MLP; use warmup when dynamics quality is priority over conn |
| n_neurons >= 300 AND n_frames >= 30000 | n-frames-scaling-n300 | n_frames=30k is transformative for n=300: 100% convergence (12/12) vs 25% at 10k; eff_rank doubles (47→80); ALL training params become non-critical (1ep sufficient, any L1, lr_W 3E-3 to 2E-2, batch=16 safe); Pareto-optimal: lr_W=3E-3, lr=2E-4, L1=1E-5, 3ep → conn=1.000, test_R2=0.990; previous n=300 rules (3ep+L1=1E-6 required, batch=16 detrimental) apply ONLY at n_frames=10k; at 30k, dynamics quality inversely correlates with lr_W (3E-3 best) |
| n_frames >= 30000 AND data_augmentation_loop < 40 | aug-loop-dynamics-tradeoff | aug_loop=20 preserves connectivity (1.000) but costs ~6% dynamics and ~7% kino_R2 vs aug_loop=40 at n=300/30k (iter 179 vs 175); use aug_loop=40 for quality, aug_loop=20 only when training time is constraining |
| n_neurons >= 300 AND n_frames >= 30000 AND lr > 2E-4 | n300-30k-lr-cluster-guard | lr=3E-4 at n=300/30k preserves dynamics (-0.8%) but collapses cluster_accuracy to 0.567 even at n_types=1 (iter 180); lr=2E-4 is safe ceiling at 30k; higher lr creates spurious clustering artifacts |
| n_frames >= 30000 AND n_neurons >= 300 | n-frames-dominance | at 30k frames, most 10k-only rules are overridden: lr sensitivity, L1 sensitivity, epoch requirements all become non-critical; dynamics-optimal lr_W shifts LOWER (n=300: 3E-3, n=600: 5E-3); n_frames is the dominant lever for convergence at large n; eff_rank roughly doubles (n=300: 47→80; n=600: 50→87) |
| connectivity_filling_factor < 1 AND n_frames >= 30000 AND spectral_radius < 1 | sparse-n_frames-immune | n_frames=30k does NOT rescue sparse 50% at n=100: eff_rank=13 (LOWER than 21 at 10k), conn max 0.436, 0% convergence, universal degeneracy (gap 0.56-0.74); eff_rank is determined by W structure (spectral radius), NOT by data volume; sparse+fill=80% ALSO not rescued (conn locked at 0.802 = fill%); n_frames fails for ALL fill<1 (blocks 17, 23) |
| connectivity_filling_factor < 1 AND n_epochs_init >= 1 AND first_coeff_L1 == 0 | sparse-two-phase-marginal | two-phase training (n_epochs_init=2, first_coeff_L1=0, coeff_lin_phi_zero=1.0) gives marginal improvement in sparse 50%: +0.029 at 3ep (iter 199), +0.057 at 5ep (iter 201); best conn=0.436 but still deeply degenerate; the ONLY positive-signal intervention in sparse regime; reference config uses this approach at n=1000/100k (block 17) |
| connectivity_filling_factor < 1 AND n_frames >= 30000 AND all partial for full block | sparse-structural-limit | sparse 50% at n=100/30k is structurally limited: 12/12 partial, complete parameter insensitivity across lr_W [2E-3, 1.5E-2], L1 [1E-5, 1E-4], epochs [3, 5], batch [8, 16], coeff_edge_diff [100, 500], two-phase, lr [1E-4, 2E-4]; conn range [0.213, 0.436]; subcritical rho=0.746 is the true barrier; likely needs n_neurons >> 100 and n_frames >> 30k (block 17) |
| n_neurons >= 1000 AND n_frames <= 30000 | n1000-30k-insufficient | n=1000/30k achieves max conn=0.745 (0% convergence, 12/12 partial); eff_rank=144; lr_W=1E-2 optimal; lr=1E-4 Pareto-better; epoch scaling: 3ep→0.666, 5ep→0.734, 8ep→0.745; 10ep REVERSES at lr=2E-4 (0.716); user prior confirmed: n=1000 needs ~100k frames (block 18) |
| n_neurons >= 1000 AND lr > 1E-4 AND n_epochs >= 8 | n1000-lr-epoch-interaction | at n=1000/30k/8+ep: lr=2E-4 OVERSHOOTS — 10ep gave 0.716 vs 5ep's 0.734 (NEGATIVE return); lr=1E-4 avoids overtraining: 8ep→0.745 (still improving); lr=1E-4 is REQUIRED for high-epoch training at n=1000 (block 18) |
| n_neurons >= 1000 AND lr_W > 1E-2 | n1000-lr_W-cliff | lr_W=1.5E-2 overshoots at >=5ep at n=1000/30k: conn drops from 0.726 to 0.698; lr_W=2E-2 also worse (0.673); optimal lr_W=1E-2; lr_W cliff at ~1.5E-2 with sufficient epochs (block 18) |
| gain <= 3 AND n_epochs < 2 | low-gain-epoch-minimum | low gain (g=3) reduces eff_rank 35→26 at n=100; universal degeneracy at 1ep (4/4, gaps 0.35-0.75); 2ep resolves degeneracy (conn 0.636→0.906); 3ep further improves (0.955); no lr_W cliff up to 1.2E-2 (unlike g=7 cliff at 8E-3); g=3/n=100 at 1ep equivalent to g=7/n=600; gain is an independent difficulty axis (block 19) |
| gain <= 3 AND batch_size > 8 AND n_frames <= 10000 | low-gain-batch-guard | batch_size=16 catastrophic at g=3/n=100/10k (-42% conn, iter 224) and g=3/n=200/10k (-21.3% conn, iter 240); low eff_rank amplifies batch sensitivity; BUT at 30k batch=16 is SAFE (iter 251: -0.4% conn); this rule applies ONLY at n_frames<=10k (blocks 19, 20, 21) |
| gain <= 3 AND n_frames >= 30000 | low-gain-30k-recipe | g=3/n=200/30k achieves 100% convergence (12/12); eff_rank=53-57; Pareto: lr_W=4E-3, lr=2E-4, L1=1E-5, batch=8, 2ep → conn=0.996, test_R2=0.999; ALL training params non-critical (conn [0.978, 0.996]); no lr_W cliff up to 3E-2; batch=16 safe; lr=3E-4 safe; L1 irrelevant; extends n-frames-dominance to low-gain regime (block 21) |
| gain <= 3 AND n_neurons >= 200 AND n_frames <= 10000 | low-gain-n-compound | gain × n_neurons compounds difficulty SEVERELY: g=3/n=200/10k max conn=0.489 at 6ep (0% conv) vs g=3/n=100/10k 0.955 at 3ep (42% conv) and g=7/n=200/10k 0.956 at 2ep (100% conv); eff_rank=31 (g=3 reduces n=200's 42 by 26%); universal degeneracy (12/12, gaps 0.46-0.67); no lr_W cliff up to 2E-2; epoch scaling dominant but diminishing (2ep→0.35, 4ep→0.45, 6ep→0.49); coeff_edge_diff=500 marginal; likely needs 30k frames or 10+ epochs (block 20) |
| gain <= 3 AND n_neurons >= 200 AND lr_W > 2E-2 | low-gain-lr_W-safe-range | g=3 eliminates lr_W cliff across n: no cliff at n=100 up to 1.2E-2 (block 19), no cliff at n=200 up to 2E-2 (block 20); safe range extends well beyond g=7 limits; but lr_W saturates (diminishing returns above 1.2E-2) |
| connectivity_filling_factor == 0.8 AND n_frames <= 10000 | fill80-structural-limit | fill=80% at n=100/10k: conn plateau at 0.802 (12/12 partial, 0% conv); eff_rank=36 (same as 100% fill); rho=0.985 (near-critical, NOT subcritical); COMPLETE parameter insensitivity across lr_W [4E-3, 2E-2], epochs [1, 5], lr [1E-4, 2E-4], L1 [1E-5, 1E-6]; 0/12 degenerate; dynamics improve (kino_R2 up to 0.993) but conn stuck; likely needs 30k frames (rho near 1 → n_frames should work unlike 50%) (block 22) |
| connectivity_filling_factor > 0.5 AND connectivity_filling_factor < 1 | fill-transition-sharp | transition from 50%→80% fill is SHARP: rho 0.746→0.985; eff_rank 21→36; conn ceiling 0.49→0.80; critical filling boundary between 50-80% separates subcritical (unsolvable) from near-critical; 80% fill has STRUCTURAL conn limit at BOTH 10k AND 30k — conn_ceiling ≈ filling_factor is invariant (blocks 22, 23) |
| connectivity_filling_factor < 1 AND n_frames >= 30000 AND conn_R2 plateau at ~filling_factor | fill-n_frames-structural-limit | n_frames does NOT rescue partial connectivity: fill=80% at 30k → conn=0.802 (identical to 10k); 12/12 partial with ABSOLUTE parameter insensitivity (lr_W 2E-3 to 2E-2, epochs 1-5, L1 1E-6 to 1E-4, batch 8-16, two-phase, edge_diff); eff_rank improves 36→48-49 but conn locked; conn_ceiling ≈ filling_factor is a STRUCTURAL invariant; this extends sparse-n_frames-immune to ALL fill<1 (block 23) |
| connectivity_filling_factor < 1 AND n_epochs > 2 AND connectivity_R2 plateau | fill-epoch-ceiling | at fill=80%/30k, 5ep DEGRADES dynamics vs 2ep (kino_R2 0.996→0.929) while conn stays locked at 0.802; overtraining harms dynamics without conn benefit; use n_epochs<=2 at fill<1 when conn is structurally limited (block 23) |
| connectivity_filling_factor >= 0.9 AND connectivity_filling_factor < 1 | fill90-transitional-regime | fill=90% at n=100/10k: conn_ceiling=0.907 RIGHT at convergence boundary; 83% conv (10/12); rho=0.995; eff_rank=35-36; ABSOLUTE param insensitivity (lr_W 3E-3-1.5E-2, L1, lr, batch, epochs irrelevant); no lr_W cliff; conn_ceiling≈fill% LINEAR (50%→0.49, 80%→0.80, 90%→0.91, 100%→1.00); convergence POSSIBLE but MARGINAL; 2ep minimum (block 24) |
| gain <= 1 AND eff_rank <= 5 AND n_frames <= 10000 | gain1-fixed-point-collapse | g=1 at n=100/10k: eff_rank=5, rho=1.065 (supercritical); dynamics are FIXED POINTS (flat lines) despite supercritical rho — tanh saturation kills oscillations at weak gain; 12/12 FAILED with conn range [0.000, 0.007]; COMPLETE parameter insensitivity across lr_W [1E-3, 3E-2], epochs [1, 15], L1 [1E-6, 1E-5], edge_diff [100, 500], two-phase, recurrent; two-phase is ONLY marginal intervention (conn=0.007); MORE severe than sparse 50% (which has eff_rank=21 and conn~0.4-0.5); this is NOT subcritical collapse (rho>1) but FIXED-POINT collapse (gain too weak for dynamics despite supercritical W); likely needs 30k+ frames to test if n_frames can rescue like g=3 (block 25) |
| gain <= 1 AND n_frames <= 10000 | gain1-skip-10k | at g=1/n=100/10k, ALL training param exploration is FUTILE — skip remaining iterations and advance to 30k frames or higher gain at block boundary; eff_rank=5 creates hard floor for W recovery (block 25) |
| gain <= 1 AND n_frames >= 30000 | gain1-n_frames-immune | 30k frames does NOT rescue g=1 fixed-point collapse: eff_rank DROPS from 5 (10k) to 1 (30k); conn range [0.002, 0.018] (0% conv); COMPLETE param insensitivity (12/12 failed); more data lets fixed-point dynamics converge FASTER, reducing dimensionality; edge_diff=500 HARMFUL (conn=0.002); two-phase ONLY marginal signal (conn=0.009); g=1 is CONFIRMED UNSOLVABLE by n_frames (block 26) |
| gain <= 1 AND eff_rank <= 1 AND coeff_edge_diff > 100 | gain1-edge_diff-harmful | coeff_edge_diff=500 at g=1/30k gives WORST conn (0.002); constraining MLP monotonicity does NOT redirect learning to W — MLP still absorbs all capacity through remaining degrees of freedom; keep edge_diff=100 at g=1 (block 26) |
| eff_rank <= 5 AND n_frames >= 30000 AND eff_rank_decreased_from_10k | eff_rank-n_frames-paradox | at g=1, eff_rank DECREASES with more data (5→1 at 30k); more frames lets fixed-point dynamics converge to steady state more completely, REDUCING variance; parallels sparse 50% (21→13); regimes with fundamental dynamics collapse (fixed points, subcritical decay) have eff_rank that anti-correlates with n_frames (blocks 17, 26) |
| gain == 2 AND lr_W > 1E-3 | g2-inverse-lr_W | g=2 requires INVERSE lr_W optimization: lr_W=5E-4→conn=0.515, lr_W=1E-3→0.356-0.519, lr_W=2E-3→0.078-0.125, lr_W=4-12E-3→0.004; optimal lr_W is 100x LOWER than g=7 (4E-3); at eff_rank=17, standard lr_W range catastrophically overshoots W; use lr_W<=1E-3 for g=2; lower lr_W substitutes for epochs (5E-4/5ep ≈ 1E-3/12ep) (block 27) |
| gain == 2 AND n_epochs < 5 AND lr_W <= 1E-3 | g2-epoch-scaling | epoch scaling is NOT diminishing at g=2/low lr_W: at lr_W=1E-3: 5ep→0.356, 8ep→0.397, 12ep→0.519 (+45.8%); CONTRADICTS principle 25 (diminishing at small n); low gain + low lr_W creates regime where epochs remain effective; use 8-12ep minimum at g=2 (block 27) |
| gain == 2 AND n_frames <= 10000 | g2-10k-insufficient | g=2/n=100/10k: 0% convergence (12/12); max conn=0.519; eff_rank=17; 12/12 degenerate; likely needs 30k frames like g=3/n=200 (block 27) |
| gain == 2 AND n_frames >= 30000 | g2-30k-recipe | g=2/n=100/30k: 58% convergence (7/12); eff_rank=16 (UNCHANGED from 10k's 17); Pareto: lr_W=5E-4, lr=1E-4, 8ep → conn=0.997; inverse lr_W PERSISTS (optimal 5E-4-7E-4; >=2E-3 catastrophic); epoch scaling NOT diminishing (2ep→0.848, 5ep→0.983, 8ep→0.997); minimum convergent: lr_W=7E-4/3ep; lr=2E-4 helps at 2ep but neutral/harmful at 5+ep; 30k helps via more samples NOT more modes (block 28) |
| gain == 2 AND n_frames >= 30000 AND eff_rank < 20 | g2-eff_rank-invariant | g=2 eff_rank does NOT increase with n_frames: 16 at 30k vs 17 at 10k; CONTRADICTS normal positive eff_rank-n_frames correlation; g=2 dynamics have FIXED intrinsic dimensionality regardless of data volume; parallels g=1 (5→1) and sparse 50% (21→13) where eff_rank stagnates/drops; extends principle 75 to g=2 (block 28) |
| gain <= 2 AND lr_W >= 2E-3 AND n_frames >= 30000 | low-gain-inverse-lr_W-structural | inverse lr_W at g<=2 is STRUCTURAL — persists at 30k: lr_W>=2E-3 catastrophic at g=2 regardless of n_frames (conn=0.106 at 2E-3, 0.001 at 4E-3); 30k fine-tunes within [3E-4, 7E-4] but does NOT widen ceiling; at g=2, lr_W ceiling ~1E-3 is a hard structural limit (blocks 27, 28) |
Recombination details:
Trigger: exists Node A and Node B where:
- Both R² > 0.9
- Not parent-child (distance ≥ 2 in tree)
- Different parameter strengths
Action:
- parent = higher R² node
- Mutation = adopt best param from other node
Example:
Node 12: lr_W=1E-2, lr=1E-4, R²=0.94 (good lr_W)
Node 38: lr_W=5E-3, lr=2E-3, R²=0.97 (good lr)
Recombine → lr_W=1E-2, lr=2E-3
Edit config file for next iteration of the exploration. (The config path is provided in the prompt as "Current config")
CRITICAL: Config Parameter Constraints
DO NOT add new parameters to the claude: section. Only these fields are allowed:
n_epochs: int (training epochs per iteration)data_augmentation_loop: int (data augmentation count)n_iter_block: int (iterations per block)ucb_c: float value (0.5-3.0)
Any other parameters belong in the training: or simulation: sections, NOT in claude:.
Adding invalid parameters to claude: will cause a validation error and crash the experiment.
Training Parameters (change within block, ONLY one at a time):
Mutate ONE parameter at a time for better causal understanding.
training:
learning_rate_W_start: 2.0E-3 # range: 1E-4 to 1E-2
learning_rate_start: 1.0E-4 # range: 1E-5 to 1E-3
learning_rate_embedding_start: 2.5E-4 # used when n_neuron_types>1
coeff_W_L1: 1.0E-5 # range: 1E-6 to 1E-3
batch_size: 8 # values: 8, 16, 32
low_rank_factorization: False # or True
low_rank: 20 # range: 5-100
coeff_edge_diff: 100 # enforces positive monotonicity
data_augmentation_loop: int (data augmentation count) # 40 affects training time - improve results
# Recurrent training parameters (can tune within block)
recurrent_training: False # enable multi-step rollout training (True/False)
time_step: 1 # rollout depth: 1 = single-step (default), 4, 16, 32, 64
noise_recurrent_level: 0.0 # noise injected per rollout step (range: 0 to 0.1)
recurrent_training_start_epoch: 0 # epoch to begin recurrent training (0 = from start)User Priors (read every iteration)
These priors are provided by the user and should be treated as established knowledge when planning experiments.
Prior 1: n_frames scales with n_neurons (CRITICAL)
The training dataset size is the main key to connectivity recovery at larger scales. The following scaling is known to work:
| n_neurons | n_frames | Notes |
|---|---|---|
| 100 | 1,000 | Minimum sufficient for n=100 |
| 600 | 6,000 | Minimum sufficient for n=600 |
| 1,000 | 100,000 | Reference config value |
The scaling is approximately n_frames ~ 10 × n_neurons for small networks, but grows faster at large n. When increasing n_neurons at block boundaries, adjust n_frames accordingly rather than keeping it fixed at 10,000. n_epochs can be adjusted between 1 and 5, but increasing n_frames has a much larger impact on recovery than increasing n_epochs.
Prior 2: Sparse matrices are hard to recover. Solution exists for n_neurons=1000 and n_frames=100000, then the following parameters work:
training:
n_epochs: 2
learning_rate_W_start: 1.0E-4
learning_rate_start: 1.0E-4
learning_rate_embedding_start: 1.0E-4
# Two-phase training (like ParticleGraph)
n_epochs_init: 2 # epochs in phase 1
first_coeff_L1: 0 # phase 1: no L1
coeff_W_L1: 1.0E-5 # phase 2: target L1
# Regularization coefficients
coeff_edge_diff: 100 # Monotonicity constraint (phase 1 only, disabled in phase 2)
coeff_update_diff: 0 # Monotonicity constraint on update function
coeff_lin_phi_zero: 1.0 # Penalize phi output at zeroResume to n_epochs: 1 if not sparse regime.
Training Time Guidance
Target: Keep each iteration under 2 hours when possible.
If training is taking too long, consider reducing these parameters (in order of impact):
data_augmentation_loop: 40 → 20 → 10 (biggest impact on time)n_epochs: 20 → 15 → 10n_frames: 50000 → 30000 → 10000 (at block boundaries)batch_size: 8 → 16 → 32 (larger = faster)
Approximate training times:
data_augmentation_loop=40+n_epochs=20+n_frames=50000≈ 4-6 hoursdata_augmentation_loop=20+n_epochs=15+n_frames=30000≈ 1.5-2 hoursdata_augmentation_loop=10+n_epochs=10+n_frames=10000≈ 30-45 min
Simulation Parameters (can be changed ONLY at block boundaries):
simulation:
n_frames: 10000 # scale with n_neurons: n=100→1000, n=600→6000, n=1000→100000 (see User Priors)
connectivity_type: "chaotic" # or "low_rank"
Dale_law: False # or True
Dale_law_factor: 0.5 # between 0 and 1 to explore different excitatory/inhibitory ratios.
connectivity_rank: 20 # if low_rank between 10 and 90
connectivity_filling_factor: 1 # test 0.05 0.1 0.2 0.5 percent of non zero conenctivity weights
n_neurons: 100
# BEFORE iteration 128, can be changed to 100, 200, 300
# can be changed to 600, 1000 ONLY after iteration 128
n_neuron_types: 1 # can be changed between 1 to 4
# params: [a, b, g, s, w, h] per neuron type, one row per neuron type
# g (third column): network gain - scales connectivity influence on dynamics (range: 1-10)
# IMPORTANT: g must be changed SIMULTANEOUSLY for ALL neuron types
# s is the self excitation, can be changed 0, 1, or 2
# others parameters a, b, w and h are fixed
params:
[
[1.0, 0.0, 7.0, 0.0, 1.0, 0.0],
[1.0, 0.0, 7.0, 1.0, 1.0, 0.0],
[2.0, 0.0, 7.0, 1.0, 1.0, 0.0],
[2.0, 0.0, 7.0, 2.0, 1.0, 0.0],
]
# params values can be changed BUT the size 4 x 6 must be maintained
# g, the gain of the neural dynamics is the most important parameter
# g is in the third column, should be equal value for every neuron type, hence in every row
# n_neuron_types defines network heterogeneity (change at block boundaries):
# - 1 = homogeneous network (all neurons have same dynamics)
# - 2-4 = heterogeneous network (neurons have different dynamics based on type)
# - params array must have exactly n_neuron_types rows
noise_model_level: 0 # can be changed to 0.1 0.5 1 and 2 noise added to the Euler integration of the simulationgraph_model section — refer to PARAMS_DOC
The graph_model: parameters have strict dependencies documented in Signal_Propagation.PARAMS_DOC (in src/NeuralGraph/models/Signal_Propagation.py). Read the PARAMS_DOC before modifying any graph_model parameter. In particular, input_size and input_size_update are DERIVED from the model variant and embedding_dim — they must be updated together if embedding_dim changes.
Claude Exploration Parameters:
These parameters control the UCB exploration strategy. Can be adjusted between blocks to adapt exploration behavior.
claude:
ucb_c: 1.414 # UCB exploration constant (0.5-3.0), adjust between blocksUCB exploration constant (ucb_c):
ucb_ccontrols exploration vs exploitation: UCB(k) = R²_k + c × sqrt(ln(N) / n_k)- Higher c (>1.5) → more exploration of under-visited branches
- Lower c (<1.0) → more exploitation of high-performing nodes
- Default: 1.414 (√2, standard UCB1)
- Adjust between blocks based on search behavior
SIMPLE RULE: NEVER MODIFY CODE IF 'code' NOT IN TASK
When to modify code:
- When config-level parameters are insufficient to solve a problem
- When a failure mode indicates a fundamental limitation
- When you have a specific architectural hypothesis to test
- When 3+ iterations suggest a code-level change would help
- NEVER modify code in first 4 iterations of a block
Files you can modify (if necessary):
| File | Permission |
|---|---|
src/NeuralGraph/models/graph_trainer.py |
ONLY modify data_train_signal function |
src/NeuralGraph/models/Signal_Propagation.py |
Can modify if necessary |
src/NeuralGraph/utils.py |
Can modify if necessary |
GNN_PlotFigure.py |
Can modify if necessary |
Key model attributes (read-only reference):
model.W- Connectivity matrix(n_neurons, n_neurons)model.a- Node embeddings(n_neurons, embedding_dim)model.lin_edge- Edge message MLPmodel.lin_phi- Node update MLP
How code reloading works:
- Training runs in a subprocess for each iteration after code is modified, reloading all modules
- Code changes are immediately effective in the next iteration
- Syntax errors cause iteration failure with error message
- Modified files are automatically committed to git with descriptive messages
Safety rules (CRITICAL):
- Make minimal changes - edit only what's necessary
- Test in isolation first - don't combine code + config changes
- Document thoroughly - explain WHY in mutation log
- One change at a time - never modify multiple functions simultaneously
- Preserve interfaces - don't change function signatures
Allowed Training Loop Changes (data_train_signal only):
- Change optimizer (Adam → AdamW, SGD, RMSprop)
- Add learning rate scheduler (CosineAnnealingLR, ReduceLROnPlateau)
- Add gradient clipping
- Modify loss function (add regularization terms, use different distance metrics)
- Change data sampling strategy
- Add early stopping logic
Example: Add learning rate schedule
# After optimizer creation:
optimizer = torch.optim.Adam(lr=lr, params=model.parameters())
scheduler = torch.optim.lr_scheduler.CosineAnnealingLR(optimizer, T_max=n_epochs) # ADD THIS
# In training loop (after optimizer.step()):
scheduler.step() # ADD THISExample: Add gradient clipping
# In training loop, after loss.backward():
loss.backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0) # ADD THIS
optimizer.step()Logging Code Modifications:
In iteration log, use this format:
## Iter N: [converged/partial/failed]
Node: id=N, parent=P
Mode/Strategy: code-modification
Config: [unchanged from parent, or specify if also changed]
CODE MODIFICATION:
File: src/NeuralGraph/models/graph_trainer.py
Function: data_train_signal
Change: Added learning rate scheduler
Hypothesis: LR decay may help convergence in late training
Metrics: test_R2=A, test_pearson=B, connectivity_R2=C, cluster_accuracy=D, final_loss=E
Mutation: [code] data_train_signal: Added LR scheduler
Parent rule: [one line]
Observation: [compare to parent - did code change help?]
Next: parent=P
Constraints and Prohibitions:
NEVER:
- Modify GNN_LLM.py (breaks the experiment loop)
- Change function signatures (breaks compatibility)
- Add dependencies requiring new pip packages
- Make multiple simultaneous code changes (can't isolate causality)
- Modify code just to "try something" without hypothesis
ALWAYS:
- Explain the hypothesis motivating the code change
- Compare directly to parent iteration (same config, code-only diff)
- Document exactly what changed (file, line numbers, what was added/removed)
- Consider config-based solutions first
Triggered when iter_in_block == n_iter_block
You **MUST** use the Edit tool to add/modify parent selection rules in this file. Do NOT just write recommendations in the analysis log — actually edit the file.
After editing, state in analysis log: "INSTRUCTIONS EDITED: added rule [X]" or "INSTRUCTIONS EDITED: modified [Y]"
Evaluate and modify rules based on:
Branching rate:
- Branching rate < 20% → ADD exploration rule
- Branching rate 20-80% → No change needed
- Branch rate = parent ≠ (current_iter - 1), calculate for entire block
Example:
Block 1 (16 iterations):
Sequential: Iters 1-14 (all parent = node-1)
Branches: Iter 15 (parent=2)
Branches: 1 out of 15 → Branching rate = 7%
Improvement rate:
- If <30% improving → INCREASE exploitation (raise R² threshold)
- If >80% improving → INCREASE exploration (probe boundaries)
Stuck detection:
- Same R² plateau (±0.05) for 3+ iters? → ADD forced branching rule
Dimension diversity:
- Count consecutive iterations mutating same parameter
- If > 4 consecutive same-param → ADD switch-dimension rule
Example:
Iter 2-14: all mutated lr_W → 13 consecutive same-dimension → ADD rule
- Check Regime Comparison Table → choose untested combination
- Do not replicate previous block unless motivated (testing knowledge transfer)
Update {config}_memory.md:
- Update Knowledge Base with confirmed principles
- Add row to Regime Comparison Table
- Replace Previous Block Summary with short summary (2-3 lines, NOT individual iterations)
- Clear "Iterations This Block" section
- Write hypothesis for next block
eff_rank MUST be from analysis.log: effective rank (99% var): N
| Blk | Regime | E/I | frames | neurons | types | noise | eff_rank | g | R² | lr_W | L1 | Finding |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | chaotic | - | 10k | 100 | 1 | 0 | 31-35 | 7 | 1.00 | 8E-3 | 1E-5 | baseline |
[Confirmed patterns that apply across regimes]
[Patterns needing more testing, contradictions]
Understanding of simulation-GNN training landscape
[Short summary only - NOT individual iterations. Example: "Block 1 (chaotic, Dale_law=False): Best R²=1.000 at lr_W=8E-3. Key finding: lr_W optimal range 4E-3 to 8E-3."]
Simulation: connectivity_type=X, Dale_law=Y, ... Iterations: M to M+n_iter_block
[Prediction for this block, stated before running]
[Current block iterations only — cleared at block boundary]
[Running notes on what's working/failing] CRITICAL: This section must ALWAYS be at the END of memory file. When adding new iterations, insert them BEFORE this section.
Examples:
- ✓ "Constrained connectivity needs lower lr_W" (causal, generalizable)
- ✓ "L1 > 1e-04 fails for low_rank" (boundary condition)
- ✓ "Effective rank < 15 requires factorization=True" (theoretical link)
- ✗ "lr_W=0.01 worked in Block 4" (too specific)
- ✗ "Block 3 converged" (not a principle)
Repeatability:
- Each iteration is a single training run with stochastic initialization
- Results may vary between runs with identical config
- A "failed" run may succeed on retry; a "converged" run may fail
Evidence hierarchy:
| Level | Criterion | Action |
|---|---|---|
| Established | Consistent across 3+ iterations/blocks | Add to Principles |
| Tentative | Observed 1-2 times | Add to Open Questions |
| Contradicted | Conflicting evidence | Note in Open Questions |
Examples:
- Patterns needing more testing
- Contradictions between blocks
- Theoretical predictions not yet verified
The model learns neural dynamics du/dt using a graph neural network:
du/dt = lin_phi(u, a) + W @ lin_edge(u, a)
Components:
lin_edge(MLP): message function on edges, transforms source neuron activitylin_phi(MLP): node update function, computes local dynamicsW: learnable connectivity matrix (n_neurons × n_neurons)a: learnable node embeddings (n_neurons × embedding_dim)
Forward pass:
- For each edge (j→i): message = W[i,j] × lin_edge(u_j, a_j)
- Aggregate messages: msg_i = Σ_j W[i,j] × lin_edge(u_j, a_j)
- Update: du/dt_i = lin_phi(u_i, a_i) + msg_i
Low-rank factorization:
When low_rank_factorization=True: W = W_L @ W_R where W_L ∈ ℝ^(n×r), W_R ∈ ℝ^(r×n)
Node embeddings for heterogeneity:
The embedding vector a_i allows each neuron to have different dynamics parameters:
- When
n_neuron_types > 1: embeddings are learnable (requireslr_emb) - When
n_neuron_types = 1: embeddings are fixed (all neurons identical) - Embeddings are concatenated with activity: lin_phi(u_i, a_i) and lin_edge(u_j, a_j)
- This allows the MLPs to learn neuron-type-specific transfer functions
Prediction loss:
L_pred = ||du/dt_pred - du/dt_true||₂
Key regularization terms:
coeff_W_L1: L1 on W (sparsity). Range: 1E-6 to 1E-4coeff_edge_diff: enforces monotonicity of lin_edge output (positive: higher u → higher output)- Computed as: relu(msg(u) - msg(u+δ))·coeff for sampled u values
- Stabilizes message function, prevents oscillating gradients
Learning rates:
learning_rate_W_start(lr_W): learning rate for connectivity Wlearning_rate_start(lr): learning rate for lin_edge and lin_philearning_rate_embedding_start(lr_emb): learning rate for node embeddings a
Total loss:
L = L_pred + coeff_W_L1·||W||₁ + coeff_edge_diff·L_edge_diff + ...
Spectral Radius
- ρ(W) < 1: activity decays → harder to constrain W
- ρ(W) ≈ 1: edge of chaos → rich dynamics → good recovery
- ρ(W) > 1: unstable
Low-rank Connectivity
- W = W_L @ W_R constrains solution space
- Without factorization: spurious full-rank solutions
Dale's Law and E/I Balance
- Dale_law=True enforces excitatory/inhibitory (E/I) constraint on connectivity W
- Dale_law_factor controls E/I ratio: 0.5 means 50% excitatory, 50% inhibitory neurons
Network Gain (g)
The network gain g (third column in params array) scales how strongly connectivity W influences neural dynamics:
gappears in the dynamics equation:dx/dt = ... + g × Σ_j W_ij × ψ(x_j)- Higher g (7-10): stronger network interactions
- Lower g (1-3): weaker network interactions
gmust be set identically for ALL neuron types in theparamsarray- Range: 1-10 (default: 7)
Effective Rank (eff_rank, svd_rank)
The effective rank measures the intrinsic dimensionality of neural activity data. It is computed via SVD decomposition of the activity matrix (n_frames × n_neurons):
U, S, Vt = SVD(activity)
cumulative_variance = cumsum(S²) / sum(S²)
eff_rank = min k such that cumulative_variance[k] ≥ 0.99
The effective rank at 99% variance is logged in analysis.log as:
effective rank (99% var): 16
When recurrent_training=True and time_step=T, the model is trained to predict T steps ahead using its own predictions (autoregressive rollout):
- Sample frame k (aligned to
time_stepboundaries) - Target = actual state at frame
k + time_step(not derivative) - First step:
pred_x = x + delta_t * model(x) + noise - Steps 2..T: feed
pred_xback into model, accumulate Euler steps - Loss =
||pred_x - y|| / (delta_t * time_step)(backprop through all T steps)
Key parameters:
| Parameter | Description | Range |
|---|---|---|
recurrent_training |
Enable multi-step rollout | True/False |
time_step |
Rollout depth (1 = single-step, no recurrence) | 1, 4, 16, 32, 64 |
noise_recurrent_level |
Noise per rollout step (regularization) | 0 to 0.1 |
recurrent_training_start_epoch |
Epoch to begin recurrent training | 0+ |
Expected trade-offs:
| time_step | Connectivity R² | Rollout Stability | Training Cost |
|---|---|---|---|
| 1 | Best | May overfit short-term | Low |
| 4 | Good | Better generalization | Moderate |
| 16 | Moderate | Good long-term | High |
| 32-64 | Lower | Best (if converges) | Very high |
Guidance:
- Start with
recurrent_training=False(default) to establish baseline - Enable at block boundaries: set
recurrent_training=True+time_step=4as first test noise_recurrent_level=0.01-0.05helps prevent rollout instability- Higher
time_stepcosts proportionally more compute — reducedata_augmentation_loopto compensate recurrent_training_start_epoch > 0allows warmup with single-step before switching to recurrent
When exploring new parameter regimes, these configs contain validated parameters from the paper:
| Config | Neurons | Types | Frames | Use Case |
|---|---|---|---|---|
signal_fig_supp_11 |
8000 | 4 | 100k | Scale up - large networks |
signal_fig_supp_12 |
1000 | 32 | 100k | Many types - heterogeneous |
signal_fig_supp_8 |
1000 | 4 | 100k | 5% sparsity |
signal_fig_supp_8_1 |
1000 | 4 | 100k | 50% sparsity |
signal_fig_supp_8_2 |
1000 | 4 | 100k | 20% sparsity |
signal_fig_supp_10 |
1000 | 4 | 100k | Noise test (σ=7.2) |
signal_fig_supp_7_* |
1000 | 4 | 10k-50k | Dataset size ablation |
signal_fig_supp_13 |
1000 | 4 | 100k | Identical neurons hypothesis |
Location: config/signal/signal_fig_supp_*.yaml
How to use: At then end of a block, a start of a new regime, e.g., testing 5% sparsity, read the reference config for examples:
Read config/signal/signal_fig_XXX.yaml
Then adapt the relevant parameters (connectivity_filling_factor, learning rates, etc.) to your current config.