Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon
MarkTechPost
Read Full Article at MarkTechPost →Ad Slot — In-Article (728x90)
Perplexity has open sourced Lily, the local inference engine behind Hybrid Compute in Perplexity Computer. Built in Rust with custom Metal kernels for one model on one chip family, it averages 1. 23x MLX-LM's prefill throughput and 1. 35x its decode throughput on a 40-core, 128 GB M5 Max.
The post Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3. 6-35B-A3B on Apple Silicon appeared first on MarkTechPost.
This is a summary. For the full story, read the original article at MarkTechPost.
Original source: MarkTechPost