Author
Feng An
Recent research
- Society & EconomicsOpen access
Contrastive representation learning for self-supervised deception detection in edge LLMs
Existing deceptive alignment detection schemes generally follow a three-step strategy: auto-labeling, supervised fine-tuning (SFT), and proximal policy optimization (PPO). In which, the detection is treated as a simple binary classification and rely on heavyweight teacher models...