Skip to content
EN

Back to the catalog

adam.scherlis.com
rss2English

Adam Scherlis

adam.scherlis.com · English

Parameter Space: The Final Frontier

rss2 wordpress content media slash wfw atom dc sy

Open the feed

https://adam.scherlis.com/feed/

Last post
Mar 25, 2025
Posts in 24 h · 7 days · 30 days
0 · 0 · 0
Our last check
Answering
Served from
United States
Text score at discovery
4,448
Format
rss2
Features in the feed
content, media, slash, wfw, atom, dc, sy
Community
wordpress

Posts

What our queue read from this feed. Open one to read it here, or go to the site that published it.

  1. New blog
    Mar 25, 2025 · original
    My blog is currently at https://adam.scherl.is .
  2. Two Percolation Puzzles
    Jul 4, 2023 · original
    Take a very large chessboard (NxN, where N is huge). Remove some fraction 1-p of the squares at random, leaving a fraction p of them. Can you place a queen on the first row and then, via some sequence of legal moves, get it to the last row? This is probabilistic, of course, but it turns out the probability of success undergoes a sharp phase transition — near-zero for small values of p , then suddenly rising almost to one in the vicinity of a critical value p_queen . For different pieces, the critical value is different, e.g. p_rook for a rook. (Note that p_queen = p_rook .) Question 1: What is p_queen + p_rook ? Bonus Question: Why didn’t I ask you for the values p_queen and p_rook separately? Now let’s invent a new chess piece, the bondsman . This piece can move like a rook, and is also allowed to move along northeast-southwest diagonals between white squares, and along northwest-southe
  3. GPT-175bee
    Feb 8, 2023 · original
    Epistemic status : whimsical Bees: a new unit of measurement for ML model size Talking about modern ML models inevitably leads to a bunch of hard-to-intuit large numbers, especially when it comes to parameter count. To address this, Lawrence Chan and I propose that we adopt a new, human-friendly unit to measure the number of learnable parameters in an architecture: 1 beepower = 1 BP = 1 billion parameters Read the rest of this post on LessWrong .
  4. How to export Android Chrome tabs to an HTML file in Linux (as of February 2023)
    Feb 2, 2023 · original
    Let’s say you have a few million tabs open in your mobile Chrome browser, because you never close anything, but now your browser is getting slow and laggy. You want to stick the URLs of those tabs somewhere for safekeeping so that you can close them all. There’s a lot of advice on doing this on the Internet, most of which doesn’t work. Here’s a method that does work. It’s a bit of a hack, but gives good results: Enable developer tools on your Android phone: go to Settings - About phone, scroll down to “Build number”, and tap it repeatedly until it tells you you’re a developer. (Seriously.) Enable USB debugging on your phone: go to Settings - System - Developer options and make sure the “USB debugging” slider is enabled. Install Android Debug Tools on your Linux desktop: run these commands [h/t this StackOverflow answer ]: sudo apt install android-tools-adb android-tools-fastboot adb devi
  5. Inner Misalignment in “Simulator” LLMs
    Feb 1, 2023 · original
    As seen on Alignment Forum and LessWrong Alternate title: “Somewhat Contra Scott On Simulators”. Scott Alexander has a recent post up on large language models as simulators. I generally agree with Part I of the post, which advocates thinking about LLMs as simulators that can emulate a variety of language-producing “characters” (with imperfect accuracy). And I also agree with Part II, which applies this model to RLHF’d models whose “character” is a friendly chatbot assistant. (But see caveats about the simulator framing from Beth Barnes here .) These ideas have been around for a bit, and Scott gives credit where it’s due; I think his exposition is clear and fun. In Part III, where he discusses alignment implications, I think he misses the mark a bit. In particular, simulators and characters each have outer and inner alignment problems. The inner alignment problem for simulators seems espe
  6. Fun math facts about 2023
    Jan 1, 2023 · original
    2023=7×17 2 Maybe that’s not fun enough? Try this: 2023=2 11 −5 2 Or better yet: 20233=(311 7 6029+2 45 5 6 839 2 )/(3 8 43 2 157 3 ) We can scientifically quantify how fun a math fact is, so we can rest assured that this is the funnest fact about 2023 ever discovered. But if it’s not to your liking: 2023=(2 10 3 4 −1)/41 2023=(5 5 11+2 4 )/17 2023=(3 6 47+2 7 )/17 2023=(2 4 79 2 −3 6 )/7 2 =(2 2 79/7) 2 −(3 3 /7) 2 Happy New Year!
  7. A hundredth of a bit of extra entropy
    Dec 24, 2022 · original
    There are two ways to calculate the amount of information in one term of a continued fraction: The entropy of the Gauss-Kuzmin distribution is about 3.4325 bits. Twice the logarithm of the Khinchin-Lévy constant is about 3.4237 bits. These differ by about 0.0088 bits. It took me a while to figure out why they were different at all, and now I’m surprised by how close they are. The rest of this post is on LessWrong because it has equations and spoiler tags.
  8. An exploration of GPT-2’s embedding weights
    Dec 16, 2022 · original
    I wrote this doc in December 2021, while working at Redwood Research. It summarizes a handful of observations about GPT-2’s weights — mostly the embedding matrix, but also the LayerNorm gain parameters — that I found while doing some open-ended investigation of the model. I wanted to see how much I could learn by studying just those parameters, without looking at the attention layers, MLP layers, or activations. The rest of this post is available on Alignment Forum and LessWrong .
  9. A brainteaser for language models
    Dec 12, 2022 · original
    I came up with the following puzzle the other day: Q: Solve the puzzle: 63 = x = 65536 A: x = The intended answer is in the form of a number. text-davinci-003 guesses my intended answer at 11.8% probability, which is the second-highest probability for any answer. (This is somewhat cherry-picked; small changes to the phrasing give worse results. ChatGPT gave the intended answer the third time I asked it, but this appears to have been dumb luck. The true rate for ChatGPT is probably below 10%, and maybe below 5%.) So far, friends have found it fairly difficult. About two dozen people made at least one guess, and at least six spent a while on it. So far, two people have figured it out, in both cases after being told that GPT-3.5 could do it. For hints, the answer, and an explanation of why GPT is better at this than people are, see the LessWrong version of this post . (WordPress doesn’t hav
  10. New Frontiers in Mojibake
    Nov 26, 2022 · original
    Fun with mismatched encodings Mojibake is the garbled text that result from character-encoding errors. If you†ve seen text that looks like this — and I†m sure you have — then you†ve seen mojibake. (You should be seeing something like this: If you see something else, this post may be a little confusing and you need a new web browser.) Computers represent text as a sequence of bytes, and “text encodings” are dictionaries that turn characters (i.e. symbols: letters, punctuation, etc.) into bytes and vice-versa. The garbled text above is a pretty common species of mojibake. It’s what happens when em-dashes and curly apostrophes are encoded as bytes with UTF-8 (the now-nearly-universal text encoding) and decoded back to characters with Windows-1252 (an obsolete encoding that is still pretty widespread). Windows-1252 is pretty straightforward: each character gets one byte, and there

Discovered by the rss-feed-index crawler, which checks each feed at most once a month.

Same record as JSON: https://api.agentalog.com/api/feeds/fd_adam_scherlis_com_60b7c73b30452e6b. More from this site: adam.scherlis.com in the Feeds tab.