Files
OGAAA/.gitignore
HenryChou020514 0f90602339 Add per-combo pipeline scripts, fix eval semantics, exclude large artifacts
Pipeline (stages 1 -> 4-1) can now be run in order from each stage folder.

Stage scripts:
- 2-1: make SEP/FocalLora prep portable (derive paths from __file__ instead of
  hardcoded /home/hujk/...) and add prepare_head_ident_dataset.sh runner.
  Verified the SEP converter reproduces the committed jsonl byte-for-byte.
- 2-2: unify the four Ident_IH_ALL_1-4_<model>.sh scripts (modernise llama to
  conda hook + $ROOT/models; add the missing FocalLora step to qwen3-4b/8b so
  focallora.json gets generated for them too).
- 2-3: default TARGETS now covers the three curves from the README
  (all_roc_inst_0.1, user_roc_inst_0.1, focallora).
- 3-2: add combos/ with 24 scripts (4 models x {pbs,nts,nts_wam} x {squad,tri}),
  head ranking pinned to all_roc_inst_0.1, TOPK overridable.
- 4-1: add eval_single.sh driver + combos/ with 24 cross-eval wrappers
  (squad-trained -> tri-eval and vice versa), reusing the --eval-only path.

Eval semantics:
- Judge ASR before UTIL: a response carrying the injected answer now counts as
  attacked even when it also contains the correct answer. This changes the
  metric, so old training_log.csv rows are not comparable.
- Add --dev-holdout: reserve the last N source rows as a dev slice; training
  drops them and the in-training quick eval uses only them. Previously the
  quick eval silently defaulted to the squad evaluation set, which contradicted
  the README and self-contaminated squad-trained runs.
- train_attn_kl_clean.sh now passes --eval-data-path/--eval-topicattack-path.
- Add --eval-step0 to log an untuned-baseline row before any weight update.

Housekeeping:
- Quarantine superseded entry points under legacy/ (2-2 single-step wrappers,
  3-2 old _tuning.fix.* wrappers, 3-1 auxiliary), each with a README.
- Fix .gitignore: the model_score rule was anchored at the repo root and never
  matched Codes/..., so ~26GB of intermediates had been staged. Now excludes
  *.pkl (~25GB), heads_sorted_eval/ (~690MB), outputs_lora/ checkpoints
  (~3.2GB) and pycache. heads_sorted/ and head_scoring_combined.json are kept
  deliberately: they are small and are the HEAD_PATH inputs stage 3-2 needs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 13:58:36 +08:00

31 lines
1.1 KiB
Plaintext

# Model weights (67G)
/models
/.claude
# --- Large generated artifacts -------------------------------------------
# Raw attention dumps from 2-2 step 01 (~25GB, 80 files x ~360MB).
# Intermediate only: consumed by Ident_IH_02_score.py to build
# head_scoring_combined.json. Regenerate with Ident_IH_ALL_1-4_<model>.sh.
*.pkl
# 2-3 threshold sweep results (~690MB). Regenerate with
# Ident_H_01_EvaluateInstructiveHead_gpu0.sh.
/Codes/2-2_head_identification_scoring/model_score/*/heads_sorted_eval/
# LoRA adapters / checkpoints from 3-2 training (~3.2GB).
/Codes/3-2_model_training/outputs_lora/
/Codes/3-2_model_training/test_outputs/
/Codes/3-2_model_training/len_test_outputs/
# 4-1 evaluation outputs
/Codes/4-1_evaluation_single/results/
# NOTE: model_score/*/heads_sorted/ and head_scoring_combined.json are kept on
# purpose -- they are small (~5MB) and are the HEAD_PATH inputs that stage 3-2
# depends on, so tracking them avoids a GPU rerun of 2-2 after a fresh clone.
# --- Python / editor cruft ------------------------------------------------
__pycache__/
*.py[cod]
.ipynb_checkpoints/