Boost Mac Speeds With iPhone AI Acceleration Tech
Overcoming Hardware Limitations
Qwen3.8-27B is a highly capable artificial intelligence model. However, it requires adequate system memory to function optimally. The M4 Pro MacBook Pro contains only 24GB of unified memory. Consequently, this hardware limitation severely restricts prefill speeds when running this AI model. According to Wccftech, one innovative user connected an iPhone 17 Pro Max to their portable Mac via USB-C. By utilizing custom software, this setup boosted the AI model’s prefill performance by up to 44 percent.
Innovative Dual-Processing Approach
To bypass this 24GB memory ceiling, Reddit user u/StayLameBro employed a clever dual-processing strategy. They intentionally split the computational workload between the local Mac and the smartphone. For prefill acceleration, the MacBook Pro processes layers 1 through 40 of every 256-token batch. Subsequently, the computer streams the activation data directly to the mobile device. Meanwhile, the iPhone utilizes its A19 Pro GPU to execute layers 41 through 64. Concurrently, the Mac begins processing the subsequent data batch immediately. Therefore, relying on GPU computation accelerates the smartphone’s processing speed by approximately 2.4 times.
Harnessing the Neural Engine
Furthermore, the Neural Engine remains highly active throughout this entire processing cycle. The system seamlessly compiles every 16K segment of older context into a dedicated Neural Engine model. At a 140K context length, single-token write times decrease significantly from 279 milliseconds to 176 milliseconds. Ultimately, this represents a vast improvement over using solely the smartphone GPU.
Analyzing Performance Enhancements
Consider the demanding task of prefilling a 2000-token document and saving the session. With the context set to 8K, the standalone Mac achieves 132 tokens per second. However, connecting the iPhone increases this processing speed to 177 tokens per second. Thus, this unique configuration delivers a substantial 35 percent performance enhancement. When expanding the context to 16K, the speed surges from 109 to 157 tokens per second. Consequently, this yields an impressive 44 percent surge in overall efficiency. Moreover, configuring the context window to 32K elevates the prefill rate from 101 to 130 tokens per second. This adjustment results in a solid 29 percent performance improvement.
Current Limitations and Future Prospects
Theoretically, connecting these flagship Apple devices yields truly exceptional AI prefill speeds. Unfortunately, as the Reddit developer pointed out, this innovative solution still presents several notable limitations. For instance, the smartphone does not accelerate text generation in scenarios below 64K. Instead, the Mac entirely handles the text generation workload independently.
Looking Ahead to the A20 Pro
Additionally, the upcoming A20 Pro chip powering the iPhone 18 Pro series promises even greater performance headroom. Primarily, this future silicon boasts vastly superior overall computational capabilities. Furthermore, it features a highly robust dual 16-core Neural Engine architecture. Reportedly, this advanced engine surpasses the internal 7-core GPU in data throughput for specialized tasks. Currently, this custom software remains open-source and available for download on GitHub. The developer appropriately named this experimental project “backburner”.
Even merely as an experiment, the tangible benefits of accelerated prefill speeds appear highly substantial. Naturally, ample room exists to further unlock this decentralized performance potential. However, this growing trend suggests future smartphones might face supply shortages similar to computer memory components.











