Getting 50 GB/S Back from the Apple Neural Engine
eiln.github.io - 101 poäng - 21 kommentarer - 273665 sekunder sedan
Kommentarer (9)
- VladVladikoff - 18704 sekunder sedanThis website hijacked my back button during a simple page load. You should fix that, it’s not an acceptable way to behave.
- bee_rider - 17315 sekunder sedanNice investigation.
It is always surprising to me when a nice round number like 1MiB results in the “bad performance” configuration (although it happens).
Are you sure erratum is the right word in this context? I usually see it used to describe the notice that a document has an error in it.
- eiln - 273665 sekunder sedanRTL performance erratum in the Apple M3 Neural Engine throttles DRAM weight streaming throughput down to 17–19 GB/s from the nominal 45–60 GB/s. Avoiding the problematic path in the kernel DMA engine's speculative prefetch ring increased Llama 3.2 1B token throughput from 10.0 to 24.3 tokens/s.
- anemll - 7703 sekunder sedanNot all systems affected M1 and M5MAX are OK
- vist_orn - 4154 sekunder sedanThe 2048-dimension resonance is striking; ruling out core contention before testing address patterns makes the eventual DMA explanation much easier to follow.
- Neywiny - 20127 sekunder sedanJust checking here- this systemverilog is a hypothetical telling of what you think is going on? Or do you have the actual source of the RTL?
- nelsonfigueroa - 7220 sekunder sedanThis goes way over my head and I don't understand most of it lol. I noticed you're still in the middle of getting your B.S. degree and you're already writing things like this...amazing.
- thenewwazoo - 16562 sekunder sedan"Apparently the memory controller's throughput has a dominant harmonic with wavelength 2048 in tensor-dimension space."
That got a laugh out of me.
- RantyDave - 17401 sekunder sedanUmmm, wow. That's really bad.
Nördnytt! 🤓