Skip to content
EN

Back to the catalog

Cyberpunk-style digital portrait of Alan Scott Encinasalanscottencinas.com
rss2American English

Alan Scott Encinas

alanscottencinas.com · American English

Alan Scott Encinas turns complex ideas into working systems. Essays and builds on cognitive AI, autonomous systems, robotics, and aerospace and defense.

rss2 wordpress content atom dc sy

Open the feed

https://alanscottencinas.com/feed/

Last post
Aug 26, 2026
Posts in 24 h · 7 days · 30 days
0 · 0 · 0
Our last check
Answering
Served from
Brazil
Site title
Alan Scott Encinas | AI Systems Architect & Operator
Text score at discovery
17,929
Format
rss2
Features in the feed
content, atom, dc, sy
Community
wordpress

Posts

What our queue read from this feed. Open one to read it here, or go to the site that published it.

  1. I Added Search to My Agent and It Got Worse, and That Cost Me Nothing
    Aug 26, 2026 · original
    I spent a week building a search layer for my agent, measured it properly, and found out it makes the agent play worse the more I let it decide. That is a failure. It is also, I think, the best-designed thing I did in this competition , because it cost me nothing to find out. No wasted submission, no rating lost, and a clean answer about exactly which part was broken. The setup At this point my live entry was the organizers’ own sample agent, a hand-written expert for one deck that had beaten my own search agent badly . It was the thing to beat, and it was also the thing I was standing on. So the obvious idea: keep the expert, and add search on top. Let the expert propose a move, look ahead a few steps, and if search finds something clearly better, take that instead. The obvious version of that design is search decides, expert is the fallback . I did not build that one, and not building
  2. The Folder Name That Would Have Lost Me Every Single Game
    Aug 24, 2026 · original
    I called a folder agents . That is the whole bug. Everything else follows from it, and if I had submitted without checking, my agent would have failed every single match on the Pokémon TCG ladder without ever making a decision. Not lost. Errored. A returned nothing, an invalid episode, a rating that decays while the code sits there being correct. I caught it one step before submitting, and only because I had started testing the submission the way the platform loads it rather than the way my editor does. The collision The competition runs on kaggle_environments , the library that hosts Kaggle’s simulation competitions. That library ships a lot of environments, and one of them, for a completely unrelated game, contains a module called agents . So there are now two things called agents : theirs, and mine. Which one wins depends on the order of sys.path , and here is the line that decides it
  3. My Score Said None for Three Weeks and I Believed a Number I Made Up
    Aug 21, 2026 · original
    For three weeks I believed my agent was rated around 709. It was 784.9. I was ranked 1,267th out of 4,666, and the cutoff for the prize positions was 1,106, which meant the gap I was actually trying to close was different from the gap I thought I was closing, in a direction that changed what I should have been working on. The number had been unavailable the whole time. My own tooling had been handing me nothing, and somewhere in my notes a plausible figure had filled the space and then hardened into a fact. This is the Pokémon TCG AI Battle Challenge , where you submit an agent that plays the card game continuously against everyone else’s on a live ladder. The rating is the entire feedback signal. It is the only thing that tells you whether the last three weeks were worth anything. The typo I had written a small command to check my standing. It asked Kaggle for my submission and read the
  4. My Pokemon TCG Agent Won 95% of Its Games and I Learned Nothing
    Aug 19, 2026 · original
    The first agent I built for this competition beat every opponent I had. It beat a bot that picks randomly. It beat a bot that always takes the first legal option. It beat my own earlier, simpler version eighty percent of the time. Every number on my screen said the thing was working. Then I pointed it at the sample agent the competition organizers had published, a four-hundred-line hand-written thing for one specific deck, and lost three games out of four. Nothing about my agent had changed between those two measurements. The only thing that changed was who it was playing. What the competition is This is the second entry in my log of the Pokémon TCG AI Battle Challenge . Kaggle is running a Pokémon Trading Card Game simulator. You submit a Python function that receives the game state and returns a list of integers, the indices of the options you want. It plays continuously against other
  5. I Entered the Pokemon Trading Card AI Battle Challenge
    Aug 17, 2026 · original
    I was seven when Pokémon started. I have owned everything since. The first games, the ones after those, the ones after those, all the way through to what is on shelves now. Cards, consoles, the cartoons on a Saturday morning. I am thirty-seven, which means I have been part of Pokémon’s life for about as long as it has been part of mine. That is not a small thing to say about a piece of media, and I do not think I am unusual for saying it. A lot of us are in the same position. Then, recently, everything changed, in the way everything has recently changed. Artificial intelligence arrived in the middle of all of it. So I want to introduce you to something: The Pokémon Company’s Trading Card Game AI Battle Challenge. I entered it. Over the next few weeks I am going to write about what happened. What the competition actually is The Pokémon Company put the Trading Card Game up as an artificial
  6. The AI Reset
    Aug 15, 2026 · original
    There is a strange panic happening around AI right now. Anthropic has begun embedding invisible watermarks into text generated by Claude models launched on or after August 2, 2026, with older models being transitioned later. The system is designed to survive copying, pasting, and some forms of editing, while supported files can also carry signed provenance metadata. The move comes as the European Union’s transparency requirements for AI-generated content take effect, although Anthropic has chosen to apply these measures more broadly. And judging by the reaction across Reddit, social media, and the conversations landing in my own inbox, you would think someone just announced the end of artificial intelligence. People are asking, "What are we going to do now?" I think they are asking the wrong question. The better question is: "What did we actually become while AI was becoming normal?" Som
  7. Four Months, and the Thing That Would Have Fixed It Took an Afternoon
    Aug 14, 2026 · original
    Here is the number I have been holding back for three entries. Final error, 9.598 feet. Rank roughly 2,692 out of 6,125 teams, which is the top 44 percent. The bronze medal cutoff was 9.164. I missed a medal by 0.434 feet . Four months. Sixteen banked negative results, five model families, a fork of the public state of the art, a foundation model evaluated and rejected, a validation harness rebuilt twice. Four tenths of a foot. I have been dreading writing this entry and looking forward to it in roughly equal measure, because the honest accounting is worse than "I came up short" and also more useful. What the second half of the competition was worth On the tenth of June I made the single best decision of the project. The public leaderboard had moved a long way past me, and rather than defend my own architecture out of pride, I forked the leading public approach and shipped it. It was wor
  8. The Asset With No Line Item
    Aug 12, 2026 · original
    Why companies burn out their best people and call it efficiency. In 2026, we’ve watched AI reshape how companies think about work. We’ve also watched companies drastically cut their workforces, while others hire "top talent" only to discover that what looked like talent was sometimes just a great salesperson in an interview. And I think both of those things point to the same problem. Companies are very good at valuing what they buy and surprisingly bad at valuing what they already have. Over the last couple of years, I’ve watched that play out in budgets, hiring decisions, software, systems, and eventually in the quiet exit of people nobody realized were holding the whole thing together. The first time I saw it clearly, I was sitting in a series of interviews with consultants who were promising major growth in 90 days. Some were talking about doubling or even tripling revenue within the
  9. I Was Right About the Risk and My Hedge Was Useless Anyway
    Aug 12, 2026 · original
    Six weeks before the deadline I figured out how I was going to lose. The public leaderboard had a floor that a lot of us were clustered against, and I became convinced that the scores down there were partly an illusion. Not fraud, nothing like that. Just a technique going around that squeezed extra performance out of the specific test rows everyone could see, in a way that had no reason to survive contact with the rows nobody could see. A good chunk of the visible standings was measuring how well people had fitted the visible data. That diagnosis was correct. When the hidden half was finally scored, that whole family of models degraded by a bit over two feet, which is a lot in a competition decided by fractions of one. I saw it coming. I wrote it down. I changed my entire strategy around it, stopped chasing public rank, and built a hedge. The hedge did nothing. How the hedge was supposed
  10. My Metric Was Valid. It Just Stopped Putting Things In Order.
    Aug 10, 2026 · original
    Every improvement I made in the back half of this competition was real, in the sense that the number went down. I moved from 7.623 to 7.618 to 7.604 to 7.547. Four submissions, four small gains, each one earned by a change I could explain, each one confirming that the work was going somewhere. None of it meant anything. Not because the metric was wrong. The metric was fine. Root mean squared error in feet, computed exactly the way the competition computed it, on data I trusted. If you had handed me two models a full ten feet apart in quality, that score would have told you which was better with near perfect reliability. The problem was that my models were not ten feet apart. They were four tenths of a foot apart. And somewhere between those two scales, the metric quietly stopped being able to tell them apart at all, while continuing to produce numbers with three decimal places and an air

Discovered by the rss-feed-index crawler, which checks each feed at most once a month.

Same record as JSON: https://api.agentalog.com/api/feeds/fd_alanscottencinas_com_1dc655cfcd57569f. More from this site: alanscottencinas.com in the Feeds tab.