Split learning is designed to keep raw webpage data on users’ devices while sending only intermediate information to a server. Researchers report that this shared information can still reveal private webpage content in language-based phishing detection systems.
Their Semantic Information Reconstruction Attack uses large language models to infer sensitive webpage details from this shared data. In tests on real-world phishing datasets, it outperformed conventional reconstruction attacks, pointing to a privacy weakness that developers may need to address.



