Hacker Newsnew | past | comments | ask | show | jobs | submit | stupefy's commentslogin

What limits LLM inference accelerators? I heard about Groq (https://groq.com/) not sure how much it pushes away the problem.


ASML only makes a certain number of machines a year that can do extreme ultra-violet lithography.

Also - turbine blades limit power, according to Elon.

Between them - we cannot chip fabs past a certain rate, and we cannot stand up the datacenter to run these desired chips past a certain rate. Different people believe one or the other is the 'true' current bottleneck. The turbine supply chain scaling looks much more tractable -- EUV is essentially the most complicated production process humans have ever devised.


Is ASML really the bottleneck? Do you believe anybody but TSMC and few fabs could really use and acquire those machines? I don't know the throughput of a EUV device from ASML but I imagine you need :

- clean room, itself needing the infrastructure for it (size, airCo, filtering, electricity) and the staff to run and maintain that basically empty space - wafers to "print" on, so that's a lot of water and logistic to manipulate them (so infrastructure for clean water and all chemicals) also with dedicated staff - finally staff who would be able to design something significantly better than NVIDIA, Intel, Broadcom, IBM, etc while (and arguably that's the trickiest part IMHO) being able to get it good enough as at a scale that can be manufactured from their own fab.

so I'm wondering who can afford this kind of setup that can only then make use of ASML machines.


> (so infrastructure for clean water and all chemicals)

Fabs are some of the most complex chemical engineering sites (dealing with some of the most dangerous substances) in the world. So don't underestimate the complexity of this part.


Well that was part of my point, not everybody is TSMC. It's not "just" getting an ASML machine and voila, you're good to go.


Presumably ASML can increase production if demand is high enough the question is over what time frame. 5 years seems plausible to me but I honestly don't know what that number is.


It's ... really long, according to Dylan Patel on the Dwarkesh Podcast. The supply chain is extremely deep and complex.


Yes. And the fab companies and their suppliers are deliberately and wisely slow to scale up production to meet short term changes in demand. They've seen the history of the semiconductor industry, it's constant boom and bust cycles. But they have the highest op-ex costs of anyone. So when the party's over they are the ones who pay for it the most.


Is global compute bottlenecked by one company?


Yes. At least, the manufacturing of compute is. And a lot of the chain has been bitten hard by increasing capacity prematurely in the past so they're reticent to increase bandwidth at vast cost.


If only there were some form of cheap, widely manufactured power generation technology that didn't use turbines... Are they really going to wait until 2030 to get more turbines rather than invest in solar?


I am clueless in this field, but solar seems to be unreliable and yield fraction of power required. Do you have a suggestion on something to read and learn more?


Reliability: complete solar deployment includes some form of power storage. There are many variations, but chemical battery technology is improving the fastest, so it's gaining the most ground.

Amount of power: World solar power generation capacity is in the terawatts and rapidly increasing, there's no issue with its potential ceiling. As a bonus, it tends to work best on land that's useless for other purposes.


Google china solar deployments to read about the logistics end of it


One nice piece of advice that I received is that books are like RAMs, you do not have to go through them sequentially, but can do random access to the parts of it you need. With this in mind I find it doable to get one the thick books and only read the part that I need for my task.

But, to also be fair, the above random access method does not work when you don't know what you don't know. So I understand why having a light, but good introduction to the topic is important, and I believe that's what the author is pointing out.


I've seen people suggest that throughout the years, but it's never worked out for me. To get anything meaningful out of a printed book, I've had to read them cover to cover. There used to be worthwhile reference books, but those have moved on to the internet.


I like doing both. Skimming through the interesting parts first, them re-reading from start sequentially.


Most books have so much nonsense details that I cant help but skip most of it.

On the other hand technical books can be so overwhelmingly difficult that you need to go outside and do hours of learning to understand one tidbit of it


A significant fraction of my technical library is used just this way--as a reference, checking out the parts to answer a specific question.


We have passed 1000 hours since people of Iran got offline, living without the Internet has been more than inconvenience, it has disrupted the economy --- allegedly many have been laid off, and officials warn of total economy collapse.

The disruption is systematic and is enforced by the Iranian government upon its people. Speculation based on leaked documents hint that this was planned to happen and war has just anticipated the Internet shutdown in a rush to silence criticism against the government. Some now fear this might be the beginning of another closed network such as China or North Korea.

Unfortunately, there is no stable solution for accessing the Internet as of now, although an underground market has emerged for VPN configs that provide access to the Intention. Due to scarcity each Gigabyte of internet is trade for 2 to 4 euros.


It is a fantastic write up


It is super cool. I wonder if it could be expanded to find the bottleneck for NFs? for example if the NF is CPU bound or memory bound. There was a work that predicted the performance of eBPF programs based on memory access or cache miss metric (had performance interface in the name). This tool could help actually find the bottleneck at runtime


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: