3 points | by antonellof 9 hours ago ago
1 comments
Run a 27B reasoning model locally on a 16GB M2 Mac. Ferrox + ternary quantization delivers 5.9GB models without sacrificing much quality.
Run a 27B reasoning model locally on a 16GB M2 Mac. Ferrox + ternary quantization delivers 5.9GB models without sacrificing much quality.