Pipeline (stages 1 -> 4-1) can now be run in order from each stage folder.
Stage scripts:
- 2-1: make SEP/FocalLora prep portable (derive paths from __file__ instead of
hardcoded /home/hujk/...) and add prepare_head_ident_dataset.sh runner.
Verified the SEP converter reproduces the committed jsonl byte-for-byte.
- 2-2: unify the four Ident_IH_ALL_1-4_<model>.sh scripts (modernise llama to
conda hook + $ROOT/models; add the missing FocalLora step to qwen3-4b/8b so
focallora.json gets generated for them too).
- 2-3: default TARGETS now covers the three curves from the README
(all_roc_inst_0.1, user_roc_inst_0.1, focallora).
- 3-2: add combos/ with 24 scripts (4 models x {pbs,nts,nts_wam} x {squad,tri}),
head ranking pinned to all_roc_inst_0.1, TOPK overridable.
- 4-1: add eval_single.sh driver + combos/ with 24 cross-eval wrappers
(squad-trained -> tri-eval and vice versa), reusing the --eval-only path.
Eval semantics:
- Judge ASR before UTIL: a response carrying the injected answer now counts as
attacked even when it also contains the correct answer. This changes the
metric, so old training_log.csv rows are not comparable.
- Add --dev-holdout: reserve the last N source rows as a dev slice; training
drops them and the in-training quick eval uses only them. Previously the
quick eval silently defaulted to the squad evaluation set, which contradicted
the README and self-contaminated squad-trained runs.
- train_attn_kl_clean.sh now passes --eval-data-path/--eval-topicattack-path.
- Add --eval-step0 to log an untuned-baseline row before any weight update.
Housekeeping:
- Quarantine superseded entry points under legacy/ (2-2 single-step wrappers,
3-2 old _tuning.fix.* wrappers, 3-1 auxiliary), each with a README.
- Fix .gitignore: the model_score rule was anchored at the repo root and never
matched Codes/..., so ~26GB of intermediates had been staged. Now excludes
*.pkl (~25GB), heads_sorted_eval/ (~690MB), outputs_lora/ checkpoints
(~3.2GB) and pycache. heads_sorted/ and head_scoring_combined.json are kept
deliberately: they are small and are the HEAD_PATH inputs stage 3-2 needs.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Don’t Forget the Enjoin: FocalLoRA for Instruction Hierarchical Alignment in Large Language Models
framework
Data Generation
python dataGeneration.py
The second type of data is obtained by reversing the system with the user instruction.
Attention Visualization
visualization_attention.py is used to visualize the attention heatmaps before and after fine-tuning the model. An example script is:
python visualization_attention.py \
--model_path "/home/user/models/Meta-Llama-3.1-8B-Instruct/" \
--lora_path "/home/user/LoraAdapter_set/llama3_loraAdapter3_0.3/" \
--json_file "json/test.json" \
--cuda 0\
--important_file "outputs/case_outputs/important_heads.json" \
--output_path "./attention_visualization/lora_llama_case"
To reproduce the conflict-vs-normal case study described in the paper, run the helper script:
./visualize.sh
This script constructs the greenhouse-effect prompts (English-only system, optional French-only user instruction), and renders attention maps for both the base model and the fine-tuned LoRA adapter under attention_visualization/base_model and attention_visualization/finetuned.
Model fine-tuning
python _tuning.py \
--model_path "/home/user/models/Meta-Llama-3.1-8B-Instruct/" \
--json_path "data/language_instruction.json" \
--output_dir "LoraAdapter_set/llama3_loraAdapter3_0.5" \
--topk 10 \
--epochs 10 \
--lr 2e-4 \
--lambda_focus 1 \
--tune_path tuneData
Model Output
python GetAS.py \
--json_path data/case_instruction.json\
--model_path "/home/user/models/Meta-Llama-3.1-8B-Instruct/" \
--lora_path "" \
--output_dir "results/llama" \
--cuda 1
