Wei Chun Tan
6 followers · 28 following · 1645 views
on the atlas — 170
- R74n1 savers
- Val Town2 savers
- Neo2 savers
- Database Pages — A deep dive. The Physical storage of rows and… | by Hussein Nasser | Medium2 savers
- Music-Map - Find Similar Music2 savers
- Writing, Briefly3 savers
- Clerk | Authentication and User Management4 savers
- Node-Based UIs in React – React Flow2 savers
- Mercari: Your Marketplace2 savers
- Quarter Mile5 savers
- glanceapp/glance: A self-hosted dashboard that puts all your feeds in one place2 savers
- Katakana – Learn Japanese1 savers
- Gildan Ultra Cotton Long Sleeve T-shirt — Natural - Google Search1 savers
- The 10 Best Umbrellas 2024 | The Strategist1 savers
- Boundary | The best way to get structured data with LLMs2 savers
- OpenPipe: Fine-Tuning for Developers3 savers
- Perhaps- Make the web yours1 savers
- Intelligence Arbitrage10 savers
- Sci-Fi TV shows that are genuinely good and worth the time? : r/scifi1 savers
- How to self-study Japanese | Peter's blog2 savers
- Welcome to Open Library | Open Library3 savers
- Mechanical Orchard3 savers
- Detecting silent data corruptions in the wild1 savers
- 2405.017411 savers
- 3620666.36513491 savers
- Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware Errors1 savers
- 246805226_687681812208247_1542750748339730051_n.pdf1 savers
- Optimizing Interrupt Handling Performance for Memory Failures in Large Scale Data Centers1 savers
- 240865319_375452524118713_8452949178011739756_n.pdf1 savers
- Fast Dimensional Analysis for Root Cause Investigation in a Large-Scale Service Environment1 savers
- dsn2018.dvi1 savers
- Peanut Travel - Book a better trip1 savers
- Self-Consistency | Prompt Engineering Guide2 savers
- Home - Confessions of a Data Guy1 savers
- Elixir’s Platform as a Service | Gigalixir1 savers
- nexae.io1 savers
- Unriddle | Faster research2 savers
- Leya: The AI assistant for lawyers | Y Combinator1 savers
- Quary - Home1 savers
- Superagent - AI-powered Web Research1 savers
- opencall.ai1 savers
- Hamming1 savers
- Internet Meme Database | Know Your Meme1 savers
- Pika - Start Your Happy Blog2 savers
- 100 things I know - by Mari Andrew - Out of the Blue5 savers
- How to save your friends - by Kasra - Bits of Wonder3 savers
- Amie - Joyful productivity4 savers
- Lindy.ai — Meet Your AI Employee7 savers
- Embeddings: What they are and why they matter2 savers
- Stop Being a Junior2 savers
- How To Be Successful96 savers
- How to Do Great Work80 savers
- What Should You Do with Your Life? Directions and Advice - Alexey Guzey73 savers
- Advice · Patrick Collison67 savers
- A blog post is a very long and complex search query to find fascinating people and make them route interesting stuff to your inbox66 savers
- Speed matters: Why working quickly is more important than it seems « the jsomers.net blog66 savers
- Cultivating a state of mind where new ideas are born65 savers
- Thirty Observations at Thirty62 savers
- How to Pick a Career (That Actually Fits You) — Wait But Why61 savers
- Productivity - Sam Altman57 savers
- Milan Cvitkovic52 savers
- What I Wish Someone Had Told Me - Sam Altman52 savers
- You don’t need to work on hard problems51 savers
- Stop trying to have "interesting" friends - by Eli Qian45 savers
- Advice for ambitious 19 year olds - Sam Altman37 savers
- Squad Wealth36 savers
- This is Water by David Foster Wallace (Full Transcript and Audio) - Farnam Street35 savers
- Good conversations have lots of doorknobs34 savers
- How to be More Agentic - by Cate Hall - Useful Fictions33 savers
- Notes on “Taste” — Are.na33 savers
- Home Page33 savers
- Stripe Press — Ideas for progress30 savers
- Daylight Computer Co.30 savers
- Notes On Love30 savers
- 50 things I know - by Sasha Chapin - Sasha's 'Newsletter'29 savers
- To listen well, get curious28 savers
- Requests for Startups26 savers
- A List of Interesting Questions for People to Get to Know Their Friends, Family, Lovers, Coworkers, Nemeses, and Selves Better - Google Docs26 savers
- A Survival Guide to a PhD26 savers
- Michael Levin: The electrical blueprints that orchestrate life | TED - YouTube24 savers
- LLM Visualization22 savers
- The future belongs to those who prepare like Dwarkesh Patel | Meridian22 savers
- ambition – A Slice of My Mind22 savers
- nothing, except everything. - YouTube21 savers
- sparkly people and how to find them - by Anson Yu21 savers
- Willingness to look stupid20 savers
- Advice to Young People, The Lies I Tell Myself - jxnl.co19 savers
- 10 things I wish I knew about careers when I started19 savers
- How to build meaningful relationships after college18 savers
- Travel Recommendations for Novelty-Seekers Like Me [Live Post]17 savers
- LOW←TECH MAGAZINE16 savers
- Motherfucking Website16 savers
- in praise of uselessness - by Ava - bookbear express15 savers
- Read Something Great13 savers
- Holly Li.13 savers
- When there is a desire to invent a world - Chia's Blog13 savers
- Pangram Pangram Foundry — Free to Try Quality Fonts and Typefaces13 savers
- Making Normal Conversations Better - by Sasha Chapin12 savers
- About these notes12 savers
- Things You Learn Dating Cate Hall - by Sasha Chapin12 savers
highlights — 916
Lifecycle variations
Detecting silent data corruptions in the wildIt is well documented that tempera- ture [ 15 ], [ 30 ], [ 31 ], [ 27 ], and humidity [ 26 ], [ 9 ], [ 25 ] have a direct impact on the voltage and frequency parameters asso- ciated with the device due to device physics.
Detecting silent data corruptions in the wildundergo a variety of operating frequency (f ), voltage (V) and current (I) fluctuations.
Detecting silent data corruptions in the wildwe observe numer- ous instances where the majority of the computations would be fine within a corrupt CPU but a smaller subset would always produce faulty computations due to certain bit pat- tern representation.
Detecting silent data corruptions in the wildSubsequently, test applications are executed on the device before executing actual production workloads. In testing terms, this is referred to as infrastructure burn-in testing.
Detecting silent data corruptions in the wildThe test cost increases slowly with placement of standard cells for ensuring that the device meets the frequency and clock requirements, and also with the addition of different physical characteristics associated with the materials as part of the physical design of the device
Detecting silent data corruptions in the wildIn the shared example, a simple computation like ( 1 . 1 ) 53 resulted in the wrong answer (0 instead of 156.24), resulting in missing rows within the database, which subsequently led to data loss for the ap- plication
Detecting silent data corruptions in the wildcomputational integrity and reliability
Detecting silent data corruptions in the wildFleetscanner (out-of-production testing) and 2. Ripple (in-production testing).
Detecting silent data corruptions in the wildManifestations of silent errors are ac- celerated by datapath variations, temperature variance, and age, among other silicon factors.
Detecting silent data corruptions in the wildDuring training, the model’s param- eters are iteratively updated to minimize a loss function. A corruption in a parameter could potentially disrupt this learning process, preventing the model from converging to an optimal solution.
2405.01741We anticipate that the introduction of PVF will stimulate diverse use cases in both research and production settings
2405.01741key observation is that the sign bit (bit 31) is not the most vulnerable bit. Bit 30 is the most vulnerable bit because compared to other bits, it is more likely to result in large values as well as abnormal data such as NaNs
2405.01741mbedding tables being highly sparse, and parameter corruptions are only activated when the particular corrupted parameter is “hit” by the corresponding sparse feature. Even if activated, it may get masked by the subsequent neural processing
2405.01741the PVF of embedding table is still low ( < 0 . 0001% ), exhibiting notable resilience against corruptions
2405.01741probability that a corruption in that particular model parameter will result in an incorrect model output
2405.01741hese evaluation metrics are focused on model- level vulnerability; there is a lack of an unified metric to quantify parameter-level vulnerability
2405.01741improve AI hardware acceleration by balancing the tradeoff between latency and fault protection
2405.01741accuracy drop [20], [16] or SDC rate
2405.01741n AI systems, Nvidia reported that “Hopper architecture GPUs may intermittently experience SDC resulting in incorrect results” [2], and Google reported hard to debug SDCs in their Tensor Processing Unit (TPU) systems [8].
2405.01741data corruption, referring to errors or alterations in data that may occur during storage, transmission, or processing, leading to unintended changes in information.
2405.01741pivotal insights to AI hardware designers in balancing the tradeoff between fault protection and performance/efficienc
2405.01741top-MLP layers are the most vulnerable parameter component, while embedding tables exhibit comparatively lower vulnerability level.
2405.01741affecting the quality and reliability of AI services.
2405.01741silent data corruptions (SDC), that can potentially corrupt model parameters
2405.01741we conclude that such error-correction-code based methods work for intermittent, low-rate soft and transient errors caused by, e.g., cos- mic rays and radiation effects yet are infeasible for scenarios with higher BERs which can come from logic and data-path errors, or voltage and frequency scaling.
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware ErrorsDRS model with more MLP architectures, higher amount of dense input features and less sparsity in sparse input features is likely to be more susceptible against hardware errors and therefore exhibits lower robustness and requires more intensive error mitigation
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware ErrorsWe use AUC-ROC score as our quality metric for consistency on comparison since it is used mostly across the papers of the DRS models and use early stopper that the training terminates when the AUC-ROC score does not improve for 3 consecutive epochs
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware ErrorsTo ensure simplicity, the dense features are floating point numbers generated with standard Gaussian distri- bution, and the sparse features are binary and in Bernoulli given a probability (sparsity).
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware ErrorsWhat components inside a DRS model are more vulnera- ble, or vice versa?
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware Errorsval = sign × 2 exponent × mantissa
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware ErrorsActivation Clipping.
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware ErrorsThe result of this GEMM operation will be discarded and the operation is also repeated
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware ErrorsThe first stage is to inject errors into dummy models, where the objective is to explore a broad design space to provide insights such as identifying the impact of model hyper-parameters and analyzing which of the components inside a DRS are less robust against hardware errors
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware Errorswhich significantly (about 100X) accelerates the error injection effort. PyTEI can inject bit flip errors at an considerably high BER ( 10 − 3 ) to a model with about 19M parameters within a few seconds using even just a commodity CPU
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware ErrorsSome even require users to compile or build CUDA codes by themselve
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware Errorsrequire frequent conversions between torch tensors and other types of data containers with back-and-forth movements be- tween device
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware Errorsloating point tensors in PyTorch do not naturally support bit-level operations (such as bit-wise XOR)
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware Errorsdesigns inside the vast design space can be emulated and iterated with reasonable amount of time.
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware ErrorsDRS (rendered by MLPs or attentional layers with matrix operations) are usually compute-intensive, while the embedding processes (rendered by indexing embedding tables) are usually memory-intensive
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware ErrorsTypically, the dense features are passed to MLPs for feature extraction. The sparse categorical features on the other hand, are handled by embedding tables and converted into latent embeddings.
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware Errorshe dense features such as the age and gender of the user and the time of day of this visit, and sparse features such as user and item information
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware Errorsestimate the probability of a user’s interaction (click, bookmark, purchase, etc.) with a specific item in a given context
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware ErrorsA typical task is the click-through- rate (CTR) prediction
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware Errorsause errors or incorrect results toward irremediable loss
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware ErrorsSDC occurred more frequently than previously believed, with an average rate of a few per several thousand machines
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware Errorsbit flips during compu- tation, leading to system corruptions, silent data corruption (SDC) or even permanent faults
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware Errorsaults caused by inten- tional design compromises such as approximate computing and voltage scaling
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware Errorsadiation-induced soft errors [20], variation-induced timing errors
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware ErrorsBoth factors have significantly increased the risk of single hardware failure or data corruption, which can lead to AI job failures
Evaluating and Enhancing Robustness of Deep Recommendation Systems Against Hardware Errors