Model Releases
Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds
arXiv:2608.13069v1 Announce Type: new Abstract: Large language models (LLMs) are predominantly aligned to function as passive, sycophantic assistants. We challenge this default paradigm by empirically
arXiv:2608.13069v1 Announce Type: new Abstract: Large language models (LLMs) are predominantly aligned to function as passive, sycophantic assistants. We challenge this default paradigm by empirically evaluating the cognitive plasticity of open-weight architectures when subjected to rigorous behavioral reprogramming. Our objective is to induce a proactive, Socratic conversational framework, characterized by high-frequency question generation under strictly constrained high-performance computing (HPC) conditions. Through a massively parallelized hyperparameter sweep comprising 405 HPC jobs, we define precise mathematical bounds for parameter-efficient fine-tuning (PEFT). We identify an architectural threshold at LoRA rank r=16 and demonstrate via extensive epoch ablation that generalization capacity strictly reaches its optimal convergence within an optimized training window of e in [2, 3] depending on dataset density (minimum validation loss of 0.919). Furthermore, scaling model capacity to 14B parameters yielded a lower localized evaluation perplexity (1.414). Subsequent Direct Preference Optimization (DPO) successfully decoupled the underlying assertive behavior from localized syntax, while rigorous cross-lingual stress testing reveals both the capabilities and the structural boundaries of zero-shot persona transfer, demonstrating robust alignment in closely related linguistic families alongside identifiable degradation pathways in morphologically distant targets. These findings establish a rigorous empirical framework for compute-efficient, cross-lingual behavioral modification.
Related
- Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization
- Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks
- Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation
- Safety Targeted Embedding Exploit via Refinement
- LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
Source: arXiv cs.AI | 2026-08-14