StartupXO
Search
English

Show XO: Performance and Usage of the Open-Type Decision AI Model Laya-MLX Running Locally on Apple Silicon (github.com)

No votes yet startupxo 1 comment

Writing language: Korean Read in the original language

Summary / Read source ↗

- Laya-MLX is a local AI model running on Apple Silicon that instantly produces decision results by choosing answers or scoring among options for input sentences. - This model operates on a proprietary MLX runtime without PyTorch or cloud support and can be easily installed and used in Python via code 공개 on GitHub. - Based on the M3 Max chip, it records median latencies of 13.4 ms for single question processing and 7.4 ms for multilingual models, with a max throughput of 146.8~395 per second when processing batches of 64 (Presented by mizorewww, Test environment: M3 Max, macOS 27.2, Python 3.12.13). - The original checkpoints and FP16 converted weights are 공개 on GitHub and Hugging Face, offering three main models suitable for English and multilingual input.
Found on

GitHub ↗ / 6,699 stars

Sign in to comment

1 comment

Editorial opinion startupxo

A notable aspect of Laya-MLX's typed decision method is that it calculates probability values directly without outputting token-level results for each question, achieving fast response times. Specifically, the performance of processing within tens of milliseconds in the M3 Max environment is supported by actual batch throughput, which adds to its reliability. However, the environmental restrictions limited to macOS 14+ and Apple Silicon might narrow its scope of use, so it's important to verify the operating system version and hardware compatibility when applying it in practice.
Writing language: Korean

Keyboard shortcuts

Choose a post with the up and down arrows, then press Enter.

↑ / ↓
Previous post / next post
Enter
Open summary and comments for the selected post
Tab
Move to the submit or comment button, then press Enter

Type normally in text fields. Tab and Enter are always available.