Aligning large language models across the lifecycle: A survey on safety-usability trade-offs from pre-training to post-training.
Where this comes from
- Record sourced from PubMed, PMID 42019220.
- Also identified by DOI 10.1016/j.neunet.2026.108996.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Large language models (LLMs) are increasingly embedded in search, productivity tools, and autonomous agents, where safety failures or degraded utility can propagate across many applications. Yet most alignment techniques are still designed and evaluated in isolation, making it difficult to see how early choices in data, objectives, and optimization interact with later fine-tuning and adaptation. This survey takes a lifecycle view of LLM alignment with the safety-usability trade-off as the organizing lens. We first examine how pre-training data curation, corpus sanitization, privacy protection, and safety-aware objectives shape baseline behavior and memorization risk. We then compare post-training alignment paradigms, including supervised fine-tuning and both RL-based and RL-free preference optimization, such as Reinforcement Learning from Human Feedback (RLHF), Reinforcement Learning from AI Feedback (RLAIF), Constitutional AI (CAI), and Direct Preference Optimization (DPO). Finally, we analyze lightweight adaptation and model editing, including parameter-efficient fine-tuning (PEFT), adapters, knowledge editing, and machine unlearning, as a second front for both eroding and restoring earlier safety guarantees. Across stages, we provide an operational safety-usability ontology, a quantitative synthesis of reported trends, and minimum evaluation checklists for resource-constrained practice. We conclude with open challenges in multimodal and cross-lingual safety, dynamic value pluralism, and providing clearer guarantees for editing and unlearning in real-world pipelines.