Sebastian J Zhao
7 followers · 11 following · 1039 views
on the atlas — 50
- pdf1 savers
- X1 savers
- Beer 101: Types and styles - The Beer Store1 savers
- 10 Types of Odd Friendships You're Probably Part Of — Wait But Why11 savers
- Explore | Listen on NTS1 savers
- KALX 90.7FM Berkeley – Ordinary people making extraordinary radio.1 savers
- KALX 90.7FM Berkeley – Ordinary people making extraordinary radio.1 savers
- linkedin/Liger-Kernel: Efficient Triton Kernels for LLM Training ·1 savers
- [2310.07831] Optimal Linear Decay Learning Rate Schedules and Further Refinements1 savers
- ESOTERIC Definition & Meaning - Merriam-Webster1 savers
- MAI-Thinking-1 | Microsoft AI1 savers
- Translations from Jianlin Su · main2 savers
- [2605.19269] CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs1 savers
- What’s MXFP4? The 4-Bit Secret Powering OpenAI’s GPT‑OSS Models on Modest Hardware2 savers
- How I Use Git Worktrees1 savers
- Ben Kamens / Claw, reflect on our work and express your feelings as animated ASCII art2 savers
- Interaction Models: A Scalable Approach to Human-AI Collaboration - Thinking Machines Lab23 savers
- [2603.18090] MOSS-TTS Technical Report1 savers
- torch.unravel_index — PyTorch 2.11 documentation1 savers
- Introducing Muse Spark: Scaling Towards Personal Superintelligence2 savers
- [2404.16710] LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding1 savers
- Kimi K2.5 Tech Blog: Visual Agentic Intelligence3 savers
- 454345238_3797399530471890_7024663226213786587_n.pdf1 savers
- [2508.16201] SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning1 savers
- [2603.23516] MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens1 savers
- [2510.11696] QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs1 savers
- [2602.17616] Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs1 savers
- [2504.13837] Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?1 savers
- LoRA Without Regret - Thinking Machines Lab37 savers
- tinker-nomics1 savers
- [2506.06105] Text-to-LoRA: Instant Transformer Adaption1 savers
- Examining the Phenomenon of Quarter-Life Crisis Through Artificial Intelligence and the Language of Twitter - PMC1 savers
- Watching the World Burn | Burning (2018) - Bright Wall/Dark Room1 savers
- 2502.110893 savers
- [2211.17192] Fast Inference from Transformers via Speculative Decoding4 savers
- 1810.048052 savers
- 2305.143142 savers
- the friendship theory of everything69 savers
- The First Fully General Computer Action Model | blog35 savers
- Making Deep Learning Go Faster29 savers
- Transformer Inference Arithmetic | kipply's blog17 savers
- Ilya 30u3012 savers
- why time felt slower when we were kids (and how to get it back)9 savers
- Generative modelling in latent space – Sander Dieleman6 savers
- it's time to start bullying young Palantir employees5 savers
- Step-by-Step Diffusion: An Elementary Tutorial5 savers
- The Bitter Lesson is coming for Tokenization | ⛰️ lucalp4 savers
- Large Language Diffusion Models3 savers
- 2409.029083 savers
- Overleaf Example2 savers
highlights — 243
Unlike ales, lagers tend to be crisp and dry
Beer 101: Types and styles - The Beer StoreThe type of yeast is what makes ales and lagers different,
Beer 101: Types and styles - The Beer Storedesigned for or understood by the specially initiated alone
ESOTERIC Definition & Meaning - Merriam-Websternder this metric, spawning more subtasks only helps if it shortens the critical path.
Kimi K2.5 Tech Blog: Visual Agentic Intelligenceuantization noise remains static throughout the process, lacking the adaptability needed to enhance exploration at critical phases
[2510.11696] QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMspromot- ing broader exploration of potential actions or tokens in RL by increasing entropy
[2510.11696] QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMsTo avoid this drawback, our analysis surprisingly reveals that quantization noise, with precise control, can benefit RL by increasing policy entropy (Fig.3).
[2510.11696] QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMsLoRA performs equivalently to FullFT for reinforcement learning even with small ranks. We find that RL requires very low capacity, a result we anticipated based on information-theoretical arguments.
LoRA Without Regret - Thinking Machines LabFor supervised fine-tuning on small-to-medium-sized instruction-tuning and reasoning datasets, LoRA performs the same as full fine-tuning.
LoRA Without Regret - Thinking Machines LabSince LoRA trains an adapter (i.e., the A and B matrices) while keeping the original weights unchanged, a single inference server can keep many adapters (different model versions) in memory and sample from them simultaneously in a batched way.Punica: Multi-Tenant LoRA Serving (Chen, Ye, et al, 2023) Modern inference engines such as vLLM and SGLang implement this feature.
LoRA Without Regret - Thinking Machines LabOn the flipside, alcohol usage is also linked to QLC.
Examining the Phenomenon of Quarter-Life Crisis Through Artificial Intelligence and the Language of Twitter - PMCWords relating to time (“night”; “weekend”; “morning”; “early”; “day”) and work (“work”; “working”) had the highest frequency and correlation strengths.
Examining the Phenomenon of Quarter-Life Crisis Through Artificial Intelligence and the Language of Twitter - PMCThe terms associated with the QLC based on the holistic model of early adult crisis are: Stuck; Trying; Leave; Change; Unemployed; Lonely; Hopeless; Overwhelmed; Unfair; Fail; Coping; Failing; Debt; Meaning; Trapped; Try; New; Identity; Sacked; Money.
Examining the Phenomenon of Quarter-Life Crisis Through Artificial Intelligence and the Language of Twitter - PMCSocial media postings that relate to actual life events and experiences can be argued to serve a developmental function, which is to represent and reify the passing of time into a simplified and publicly documented life story that can help the individual create a meaningful ongoing narrative of how their life is changing (Rettberg, 2009).
Examining the Phenomenon of Quarter-Life Crisis Through Artificial Intelligence and the Language of Twitter - PMClocked out of adult commitments (being unable to find work or love), or the feeling of being locked in to life roles that are then experienced as a poor fit for one’s identity,
Examining the Phenomenon of Quarter-Life Crisis Through Artificial Intelligence and the Language of Twitter - PMCEarly adult crisis episodes typically occur toward the latter end of the life stage of emerging adulthood, and last approximately a year
Examining the Phenomenon of Quarter-Life Crisis Through Artificial Intelligence and the Language of Twitter - PMCEpidemiological data shows that most young adults in Western countries now choose to wait for a decade or more after turning 18 before having children, or before starting a marriage or civil partnershi
Examining the Phenomenon of Quarter-Life Crisis Through Artificial Intelligence and the Language of Twitter - PMCUsers who refer to a QLC were found to post more about feeling mixed emotions, feeling stuck, wanting change, career, illness, school, and family.
Examining the Phenomenon of Quarter-Life Crisis Through Artificial Intelligence and the Language of Twitter - PMCIn losing Hae-mi and murdering Ben, Jong-su undoubtedly suffers, but in many respects, he is the author of his own pain.
Watching the World Burn | Burning (2018) - Bright Wall/Dark RoomYes, Hae-mi did not disappear or die by his hand, but considering Jong-su’s many reflexive cruelties throughout the film, she died numerous little deaths unbeknownst to him.
Watching the World Burn | Burning (2018) - Bright Wall/Dark Roomyou construct an angry jeremiad out of your nostalgia for the props of the old reality and the architecture that's been demolished.
In Response To Jean Baudrillard (Hayles, Porush, Landon, Sobchack, Ballard)gifts" of imagination and transcendence, but enhance them
In Response To Jean Baudrillard (Hayles, Porush, Landon, Sobchack, Ballard)The borders separating simulations from reality are important because they remind us of the limits that make dreams of technological transcendence dangerous fantasies.
In Response To Jean Baudrillard (Hayles, Porush, Landon, Sobchack, Ballard)Even within the boundaries of simulations, material intractability often breaks in.
In Response To Jean Baudrillard (Hayles, Porush, Landon, Sobchack, Ballard)Only when these boundaries do not exist, or cease to signify that one has left the simulation and entered reality, does the dreamscape that Baudrillard evokes shimmer into existence.
In Response To Jean Baudrillard (Hayles, Porush, Landon, Sobchack, Ballard)The Iowa farmer who has spent the day inspecting his seed corn, feeding his hogs, and spreading manure on his garden will not be easily persuaded that he lives in a world where it is no longer possible to distinguish between simulation and reality.
In Response To Jean Baudrillard (Hayles, Porush, Landon, Sobchack, Ballard)They speak of the pursuit of “social justice” when what they really mean is abolishing free markets. They speak of “equity” when what they really mean is illegal and highly unpopular racial quotas.
What Does Postmodernism Really Amount To? - Claremont Review of BooksWe’ve seen now in soaring crime rates the danger of indiscriminate hostility to all existing customs and dogmas.
What Does Postmodernism Really Amount To? - Claremont Review of Booksit’s hard to find any significant original thought in postmodernism.
What Does Postmodernism Really Amount To? - Claremont Review of BooksDerrida and his followers completely misunderstood what Saussure was saying here.
What Does Postmodernism Really Amount To? - Claremont Review of Booksfor instance, maybe em-dashes just read more conversational, so they were preferred by RLHF-ers
Why do AI models use so many em-dashes?if em-dashes are common because they’re a feature of late-1800s/early-1900s writing, why doesn’t AI prose read more like Moby-Dick?
Why do AI models use so many em-dashes?State-of-the-art models rely on late-1800s and early-1900s print books for high-quality training data, and those books use ~30% more em-dashes than contemporary English prose.
Why do AI models use so many em-dashes?multi-GPU inference sig- nificantly increases the hardware cost of deploying MoEs.
MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensatorswhat they actually mean is that they cannot imagine slumming their twenties on less than six figures a year!
it's time to start bullying young Palantir employeesBut it’s worth noting that few low-income students statistically go into fields like consulting.
it's time to start bullying young Palantir employeesYour ancestors’ wildest dreams was not to work at lockheed martin stop lying on them folks.
it's time to start bullying young Palantir employeesBut it’s also just a feature of being a mortal on earth.
it's time to start bullying young Palantir employeesfear of post-grad risk and uncertainty — incidentally, the very conditions in which bad bitches are made
it's time to start bullying young Palantir employeesHer co-founder is a world-famous climate activist who once publicly partook in a Black Friday strike with Jane Fonda against consumerism in 2019. Again, you cannot make this up!)
it's time to start bullying young Palantir employeesneoliberalism, the bitch-ass belief
it's time to start bullying young Palantir employees“Fifteen years earlier, graduates of the country’s most elite colleges had often been concerned with trying to improve the state of the world. Now, the focus was different: How can I be as financially successful as possible?”
it's time to start bullying young Palantir employees"We think what happened is that the cats sort of domesticated themselves,"
A Brief History of House CatsAll domestic cats, the authors declared, descended from a Middle Eastern wildcat, Felis sylvestris, which literally means "cat of the woods." Cats were first domesticated in the Near East, and some of the study authors speculate that the process began up to 12,000 years ago.
A Brief History of House CatsDuring the generation process, as the context continually evolves and may include mispredicted tokens, the confidence of otherwise stable tokens can fluctuate or even regres
CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace CreditsFor each sequence in the batch, if its minimum score within the current window min( S W j ) falls below a predefined threshold τ thresh , a refinement cycle is initiated for its window W j
Review, Remask, Refine (R3): Process-Guided Block Diffusion for Text Generationather than exploring uniformly, we target medium-confidence positions near an information level c info ≈ 0 . 2 .
From Bits to Rounds: Parallel Decoding with Exploration for Diffusion Language Modelshen a previous block lacks high-confidence tokens (none exceed threshold C ), ETE unmasks the highest-confidence masked token in that block
From Bits to Rounds: Parallel Decoding with Exploration for Diffusion Language ModelsETE allows high-confidence tokens in earlier blocks to be unmasked in parallel with the current block
From Bits to Rounds: Parallel Decoding with Exploration for Diffusion Language Modelse assign a uniform sampling budget of N diffusion iterations per block and unlock the next block once this budget is exhausted. This prevents the algorithm from stalling in low-throughput regimes where blocks are nearly complete and few tokens remain to be decoded
From Bits to Rounds: Parallel Decoding with Exploration for Diffusion Language Models