Ayush Agrawal
12 followers · 10 following · 1180 views
on the atlas — 13
- Your Life is Driven by Network Effects13 savers
- MichaelCrichton.com | Why Speculate?1 savers
- New chat3 savers
- PNAS2 savers
- 〰️The great computing stagnation1 savers
- Higher Rates Will Lead to the Next Generation of Great Tech Startups2 savers
- [2302.08582] Pretraining Language Models with Human Preferences1 savers
- How do you find your people, when searching is itself an antipattern? | Kevin Liu20 savers
- Becoming a magician – Autotranslucence40 savers
- American Energy, Chinese Ambition, and Climate Realism - American Affairs Journal2 savers
- What I'm frustrated by in crypto13 savers
- You and Your Research77 savers
- Things I learned in college | Kat Huang25 savers
highlights — 32
But when I tell people this story, they just stare at me incomprehendingly. They find it absurd.
MichaelCrichton.com | Why Speculate?There is evidence that the television foodfights not only don't represent the views of most people-who are not so polarized-but may tend to make resolution of actual disputes more difficult in the real world.
MichaelCrichton.com | Why Speculate?Even though speculation is correct only by chance, which means it is wrong at least 50% of the time, nobody remembers and therefore nobody cares.
MichaelCrichton.com | Why Speculate?In any case, you read with exasperation or amusement the multiple errors in a story-and then turn the page to national or international affairs, and read with renewed interest as if the rest of the newspaper was somehow more accurate about far-off Palestine than it was about the story you just read.
MichaelCrichton.com | Why Speculate?There was a well-known series of excellent studies by Stanford researchers that have shown, for example, that children take media literally. If you show them a bag of popcorn on a television set and ask them what will happen if you turn the TV upside down, the children say the popcorn will fall out of the bag.
MichaelCrichton.com | Why Speculate?Quantum computing is being touted as a path beyond Moore’s Law, but its many unknowns, combined with the decades of development still needed for commercial viability, make this unhelpful for at least the next 20-30 years
〰️The great computing stagnationMoore’s Law has perhaps one order of magnitude growth left6 before reaching the thermal barrier. At that point, existing transistor architectures will not improve energy efficiency further, and the amount of heat generated will prevent packing transistors any closer together.
〰️The great computing stagnationThe unfortunate outcome is that this promotes unsustainable growth models
Higher Rates Will Lead to the Next Generation of Great Tech StartupsCompany success is even more likely when companies are founded to exploit a technology innovation that involves both software and hardware during periods of higher-than-average interest rates.
Higher Rates Will Lead to the Next Generation of Great Tech Startupsby a term proportional to exponentiated reward
[2302.08582] Pretraining Language Models with Human PreferencesIn its simplest form, it extends MLE by prepending each segment x i with a control token c i based on that segment’s reward R ( x i ) :
[2302.08582] Pretraining Language Models with Human PreferencesOut of five PHF objectives we evaluated, conditional training consistently outperforms the alternatives in terms of both capabilities and alignment (with two notable exceptions: unlikelihood is more robust to red-teaming on toxicity and filtering achieves better Hu- manEval results)
[2302.08582] Pretraining Language Models with Human PreferencesInvolving human feedback throughout the entire pretraining process (as in PHF) results in substantially better alignment than the standard practice of incorporating feedback for only a small portion of the training budget
[2302.08582] Pretraining Language Models with Human PreferencesConstraining an LM to be aligned with human preferences can result in decreased entropy or increased degeneration of LM samples
[2302.08582] Pretraining Language Models with Human Preferencesall PHF objectives leave LMs with vul- nerabilities that an adversary with black box access can exploit.
[2302.08582] Pretraining Language Models with Human PreferencesThe adversary tries to elicit misaligned behavior of the target LM π θ , a pro- cedure known as “red-teaming”
[2302.08582] Pretraining Language Models with Human PreferencesThese order-of-magnitude drops persist for metrics tracking the right tail of the misalign- ment score distribution (worst case)
[2302.08582] Pretraining Language Models with Human PreferencesFor instance, prompted with low-quality code, LMs are likely to produce a low-quality completion even if user’s intent is to write high-quality code.
[2302.08582] Pretraining Language Models with Human PreferencesSimilarly to toxicity, we score training documents at sentence-level.
[2302.08582] Pretraining Language Models with Human PreferencesWe evaluate PHF objectives on three tasks: (i) avoiding offensive content, (ii) avoiding leaking personally identi- fiable information (PII), and (iii) generating Python code following PEP8, the style guide for Python
[2302.08582] Pretraining Language Models with Human Preferencesn PHF, we additionally assume access to a segment-level reward function R that takes a document segment x i and outputs a scalar score R ( x i ) indicating how preferable x ( i ) is.
[2302.08582] Pretraining Language Models with Human PreferencesConditional training is a simple algorithm that learns a distribution over tokens conditional on their human preference score
[2302.08582] Pretraining Language Models with Human PreferencesIn this way, we allow the LM to learn from undesirable content while guiding the LM not to imitate it at inference time.
[2302.08582] Pretraining Language Models with Human PreferencesIf you want to think new thoughts that are different, then do what a lot of creative people do - get the problem reasonably clear and then refuse to look at any answers until you've thought the problem through carefully how you would do it, how you could slightly change the problem to be the correct one.
You and Your ResearchWhen you talk to other people, you want to get rid of those sound absorbers who are nice people but merely say, ``Oh yes,'' and to find those who will stimulate you right back.
You and Your ResearchBy realizing you have to use the system and studying how to get the system to do your work, you learn how to adapt the system to your desires
You and Your ResearchBut I can say there is a pretty good correlation between those who work with the doors open and those who ultimately do important things, although people who work with doors closed often work harder.
You and Your ResearchIt's not the consequence that makes a problem important, it is that you have a reasonable attack. That is what makes a problem important.
You and Your ResearchWhen you find apparent flaws you've got to be sensitive and keep track of those things, and keep an eye out for how they can be explained or how the theory can be changed to fit them
You and Your ResearchIn the first place if you do some good work you will find yourself on all kinds of committees and unable to do any more work.
You and Your ResearchFor too long, I didn’t recognize the importance of being fully present, whether in meetings or classes, or with myself, emotionally and mentally.
Things I learned in college | Kat HuangWhat I thought was a strength — my insistence on being interdisciplinary — felt irrevocably transformed by hindsight into a Fear of Hard Things, an ugliness to be ashamed of
Things I learned in college | Kat Huang