Roshan Swaroop
2 followers · 5 following · 282 views
on the atlas — 8
- How to be More Agentic - by Cate Hall - Useful Fictions26 savers
- Business Expertise - Commoncog2 savers
- Will Scaling Solve Robotics?: Perspectives From Corl 2023 | Nishanth J. Kumar3 savers
- How the most successful B2B startups came up with their original idea8 savers
- Papers I’ve read this week, Mixture of Experts edition2 savers
- Efficient LLM inference - by Finbarr Timbers3 savers
- How we reduced the cost of building Twitter at Twitter-scale by 100x – Blog1 savers
- Founder’s Lessons: Phoebe Yao, CEO of Pareto | by Taylor Fang | Foothill Ventures | Medium3 savers
highlights — 7
Given this interrelationship between the three factors, you may visualise business expertise as a triad:
Business Expertise - CommoncogTwo friends and I maniacally studied reads together, and we all had out-of-distribution results. But when we’d tell other pros what we were doing, the response from most was “nuh-uh, that’s not a thing.” They weren’t willing to consider the possibility that reads were valuable, maybe because they didn’t want to feel obligated to study them. All of my agency hacks are kind of like this, in my opinion -- big, glaring edges that people might rather ignore.
How to be More Agentic - by Cate Hall - Useful FictionsMore than half of the people we talked to just started cursing, unprompted. Two people voluntarily told me, ‘I use [competitor name], and my password is fuck[competitor name].’”
How the most successful B2B startups came up with their original ideaThe basic idea behind GPTQ is that, while there’s necessarily a drop in information contained within the network by reducing the number of bits, we can reduce the impact it has on inference accuracy by training weights to directly minimize the difference between the two:
Efficient LLM inference - by Finbarr TimbersIn my opinion, the literature indicates a clear & obvious ranking: distillation is strictly better than training a smaller model, and quantizing is probably better than training a smaller model.
Efficient LLM inference - by Finbarr TimbersYou can quantize the parameters of your model (quantization), where you keep your model exactly the same, but use less precision for each of the parameters. You can distill a smaller version of your model (distillation), where you copy the architecture of your model to make it smaller and/or more efficient and then train this new, smaller model to mimic the outputs of the original, large model. You can spend a bunch of time profiling your code and reduce the overhead without changing the architecture or parameters (optimization).
Efficient LLM inference - by Finbarr TimbersThe Diamond Age by Neal Stephenson has been really influential in my startup journey and pretty much inspired Pareto’s social impact mission.
Founder’s Lessons: Phoebe Yao, CEO of Pareto | by Taylor Fang | Foothill Ventures | Medium