Publication

A Survey on Split Learning for LLM Fine-Tuning: Models, Systems, and Privacy Optimizations

Apr 27, 2026 · 5 authors · 50 topics

Authors

Zihan LiuYizhen WangRui WangXiu TangSai Wu

Topics

Privacy-Preserving Technologies in DataAdversarial Robustness in Machine LearningBig Data and Digital Economy1 Introduction Large Language Models (LLMs) such as GPT [8, 115], Qwen [154], and DeepSeek [84] have demonstrated exceptional capabilities in general-purpose tasks. To adapt these models to specific domains, fine-tuning has emerged as the mainstream paradigm, effectively transitioning them from generalists to specialized domain experts. However, withManuscript submitted to ACM 1 the increasing scale of model parameters and the gradual depletion of high-quality public datasets, fine-tuning on private data faces the dual challenges of severe resource constraints and privacy leakage risks. Modern LLMs typically comprise tens to hundreds of billions of parameters [100], rendering full fine-tuning prohibitively expensive for small and medium-sized enterprises (SMEs) with limited computational budgets. Although Parameter-Efficient Fine-Tuning (PEFT) techniques mitigate these costs by updating only a small subset of parameters, the substantial GPU memory requirements for loading and serving large models remain a significant barrier. Furthermore, fine-tuning on proprietary data raises privacy concerns regarding data outsourcing: enterprises are often hesitant to upload sensitive data to third-party cloud servers due to risks of direct exposure and potential breaches. To address these key bottlenecks, split learning (SL) [42, 134] has emerged as a promising distributed learning paradigm. In SL, the model is partitioned between clients and a central server: clients retain and execute only a small portion of the model (typically the front-end and/or back-end model layers), while the computationally intensive middle layers are offloaded to the server. Training proceeds via the exchange of intermediate representations (smashed data) and gradients rather than raw data, thereby significantly reducing local computational and storage burdens while preserving data privacy. Recently, combining SL with LLM fine-tuning has attracted growing research attention [12, 76, 81, 91]. By applying this split-based paradigm to large language models, resource-constrained clients can participate in fine-tuning billion-parameter models while maintaining only a small fraction of parameters locally. This approach enables small and medium-sized enterprises and research institutions to collaboratively train or fine-tune private large models without exposing sensitive data to external parties, effectively addressing both resource and privacy challenges simultaneously. Integrating split learning with LLM fine-tuning (split-based LLM-FT) has enormous potential and has spurred a wealth of research across three complementary directions. First, a substantial line of work targets model-level optimization challenges in multi-client federated scenarios, addressing convergence degradation caused by biased intermediate features due to data heterogeneity (Non-IID), server-client update imbalance arising from increased server-side batch sizes when aggregating smashed data, and catastrophic forgetting during sequential training or partial client participation. Second, another research stream focuses on system-level efficiency improvements, tackling severe communication bottlenecks from frequent transmission of massive intermediate activations and gradients, resource utilization imbalances caused by server-client coordination overhead, and the straggler effect induced by heterogeneous client computing power and network bandwidth. Third, a growing body of literature investigates privacy and security concerns, encompassing both attack methodologies that exploit semantically rich gradients and smashed data to reconstruct private information, as well as the distributed framework’s exposed model manipulation interfaces to inject malicious behaviors, and their corresponding defense strategies. Several surveys on split learning have already been published, as summarized in Table 1. For instance, [52] cate gorizes SL paradigms into model-only segmentation, weight-based aggregation, and intermediate data aggregation methods, evaluating their performance under IID and non-IID scenarios. [78] surveys resource-efficient frameworks and optimization strategies for alleviating computational and communication bottlenecks caused by stragglers. [62, 120] summarize attack and defense approaches in split learning, while [41] evaluates perturbation-based and learning-based defenses against data reconstruction and label inference attacks and their impact on downstream performance. However, despite these foundational efforts, existing surveys suffer from three critical gaps in the context of LLM adaptation. First, current surveys primarily focus on traditional, small-scale models, overlooking the distinct challenges of integrating split learning with LLMs. From a model perspective, they assess optimization strategies in conventional tasks without verifying their efficacy for LLM fine-tuning. At the system level, they neglect the unique architectural demands, massive parameter scales, and performance constraints of LLMs. Regarding privacy, they overlook the Manuscript submitted to ACM| A | Survey | | on | Split Learning | | for LLM | Fine-Tuning: | Models, | Systems, | and Privacy Optimizations | | | | 3 | || --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- || | | | | | Table | 1. Comparison | of | related surveys | and research | works of split learning. | | | | | || | Paper | | Year | | | | Focus Areas | | | Unified Training | Component-Level | | | | || | | | | | | | | | | | Bottleneck | LLM | | | || | | | | Model | Performance | | System Optimization | | Privacy | Pipeline | | | | | || | [24] | | 2022 | | ✓ | | ✓ | | ✓ | | | | | | || | [78] | | 2024 | | | | ✓ | | | | | | | | || | [53] | | 2025 | | ✓ | | ✓ | | | | | | | | || | [117] | | 2025 | | ✓ | | ✓ | | | | | | | | || | [52] | | 2025 | | ✓ | | ✓ | | | | | | | | || | [56] | | 2025 | | ✓ | | ✓ | | ✓ | | | | | | || | [120] | | 2025 | | | | | | ✓ | | | | | | || | [41] | | 2025 | | | | | | ✓ | | | ✓ | | | || | [62] | | 2025 | | | | | | ✓ | | | | | | || | Ours | | - | | ✓ | | ✓ | | ✓ | ✓ | ✓ | ✓ | | | || exacerbated | | | | vulnerabilities | | of auto-regressive | | LLMs, where | training | labels act as shifted | input sequences, making | label | | | || leakage | | | mathematically | | | equivalent | to exposing | raw proprietary | | text. Second, previous | surveys treat split | learning | | | || paradigms | | | as | macro-level | | architectures | without | establishing | a unified | framework for | split-based LLM fine-tuning. | | | | || They | | lack | an | explicit, | component-level | | decomposition | | of the computational | and communication | steps involved | | | in | || split-based | | | | LLM fine-tuning, | | leaving | the training | pipeline | ambiguous. | Third, current | surveys do not systematically | | | | || map | | multidimensional | | | challenges | and | existing | solutions | to the specific | execution stages | of LLM adaptation. Without | | | a | || unified | | training | | pipeline, | | there is no | structured | way to anchor | inefficiencies, | model degradations, | privacy threats, | | and | | || their | | corresponding | | | countermeasures | | to precise | operational | components. | This omission | hinders the development | | of | a | |exacerbated vulnerabilities of auto-regressive LLMs, where training labels act as shifted input sequences, making label leakage mathematically equivalent to exposing raw proprietary text. Second, previous surveys treat split learning paradigms as macro-level architectures without establishing a unified framework for split-based LLM fine-tuning. They lack an explicit, component-level decomposition of the computational and communication steps involved in split-based LLM fine-tuning, leaving the training pipeline ambiguous. Third, current surveys do not systematically map multidimensional challenges and existing solutions to the specific execution stages of LLM adaptation. Without a unified training pipeline, there is no structured way to anchor inefficiencies, model degradations, privacy threats, and their corresponding countermeasures to precise operational components. This omission hinders the development of a comprehensive taxonomy needed to guide component-level design and future research directions. To bridge this critical gap, this paper presents the first comprehensive review focused exclusively on the intersection of split learning and LLM fine-tuning. Our primary contributions are threefold:• A unified split-based LLM-FT pipeline: Diverging from previous coarse classifications, we abstract and decompose existing methods into a unified, highly granular training pipeline. This structural breakdown enables precise analysis of component-level bottlenecks and future optimization trajectories.• multi-dimensional bottleneck analysis: We map critical vulnerabilities directly to specific execution steps, systematically exposing the root causes of convergence degradation, system inefficiencies, and privacy threats in split-based LLM-FT.• component-anchored literature taxonomy: By mapping multidimensional challenges and state-of-the-art opti mization methods directly onto our unified pipeline, we deliver a comprehensive taxonomy. This framework explicitly anchors system inefficiencies, model degradations, and privacy countermeasures to specific execution steps, systematically guiding future research. The rest of the article is structured as follows and the structure is in Figure 1. Section 2 provides a comprehensive overview of split-based LLM-FT, detailing model partitioning strategies, overarching frameworks, and our proposed unified pipeline components. Section 3 systematically analyzes strategies for model-level optimization and performance enhancement. Section 4 explores current methodologies for system-level optimization. Section 5 delves into the privacy and security landscape, categorizing prominent attack vectors and their corresponding defense mechanisms within split-based architectures. Finally, Section 6 concludes this survey.Manuscript submitted to ACMSplit-based LLM-FT § Introduction§ Overview of Split-based LLM-FT § System-Level Optimizations § Model-Level Optimizations § Privacy and Security § Conclusion§ Model Partitioning Strategies § Split-based LLM-FT Framework Variants § Split-based Pipeline Components§ Overcoming Catastrophic Forgetting § Mitigating Data Heterogeneity (Non-IID)§ Bridging Server-Client Update Imbalance § Split Point Selection § Computation scheduling and memory managing § Parameter-efficient fine-tuning (PEFT)§§ Communication Overhead Reduction § Activation and gradient compression § Communication frequency reduction § Model Aggregation Strategies § Server-Client Collaborative Optimization § Gradient aggregation.§ Asynchronous training strategies § Global Joint Optimization § Attacks on Split-based LLM-FT § Defenses on Split-based LLM-FT§ Smashed data inversion attacks § Gradient-based label reconstruction attacks § Bidirectional reconstruction attacks § Membership inference attacks § Property inference attacks § Data poisoning attacks § Backdoor attacks§ Perturbation-based defenses § Learning-based defenses § Detection-based defensesDecoupling personalization and generalization Smashed data aggregation and scheduling§ §§ Client clustering and pairing § Smashed data mixup § Learning rate scaling § Knowledge retention.Server-side memory management and computation scheduling§ Parallel training strategies § Mapping Pipeline Components to Core BottlenecksFig.

About

PublishedApr 27, 2026
TypePreprint
Citations0

Powered by the Exa API