18 comments

  • nullbio 16 hours ago
    It seems to me this is only useful for self-improvement up to the point of accomplishing objectives and problems that humans have already clearly defined, solved and mapped. For example, "At checkpoints, a stronger evaluator can replace the old one if it performs better on trusted ground-truth examples." - if they're a trusted ground-truth, they must be rigorous. If we're attempting to solve problems that humans have not already solved, where do you get your ground-truth examples? It's not like the AI is going to be able to generate these for you if it has never seen a solution to the problem.

    I can definitely see the argument that this allows us to train models faster and converge faster, because if you scale difficulty of evaluation alongside the learners capabilities, it spends a lot less time floudering around. It basically works out to be loss minimization through strategic ordering of the training data. Is that the goal here though? Or is the goal recursive self-improvement and solving problems that are currently outside of reach? Because it doesn't feel like the latter would be possible with this design.

    For example, how do you quantify "the evaluation gets -harder- as the agent gets better". Harder, how? In what direction? Via what criteria or measure?

    • AlexAndreiIacob 14 hours ago
      For tasks with no existing ground-truth data, one could try to use an adversarial objective on its own by making AI-generated outputs compete against each other, although this is beyond the scope of our experiments. Adapting our method to this setting is quite straightforward.

      For now, we have looked into composite objectives trading off performance on the ground truth against the ability to reject generated samples produced in earlier epochs. For example, when using this system to co-evolve paper writers and reviewers, we reuse generated papers that were accepted by the reviewer of epoch t as adversarial samples in epoch t+1. Then, new reviewers are rewarded for rejecting AI-generated papers. You can have ground-truth performance account for x of the utility of a reviewer, with the adversarial objective then providing 1-x.

  • robotresearcher 21 hours ago
    Here’s a paper by Floreano at EPFL from 1997 explicitly on Red Queen dynamics for creating neural networks for intelligent robot control.

    There was lots of discussion of these ideas in the 1990s. In those days we trained very small NNs - tens of nodes - by evolving their weights and topologies. A run could take days on a workstation of the time.

    This particular paper is about co-evolving predator and prey, where the behavior of each is the ‘evaluation’ of the other.

    https://infoscience.epfl.ch/entities/publication/a65d0679-68...

    • grumbelbart2 12 hours ago
      Funnily enough, Jürgen Schmidthuber invented / coined the term "Gödel Maschine" in 2003: https://arxiv.org/abs/cs/0309048

      He'll probably have a field day over this.

      • AlexAndreiIacob 11 hours ago
        For broader context: our work directly builds upon Jürgen Schmidhuber’s Huxley-Gödel Machine, while our research has been cited in a recent survey he co-authored: https://arxiv.org/abs/2607.13104

        Anyone interested in this field should definitely engage thoroughly with his body of work.

        • grumbelbart2 10 hours ago
          I by no means wanted to suggest the opposite (that's how I found the connection)!

          Schmidthuber is famous for his groundbreaking work in machine learning, but also for being somewhat left out when people list the "Grandparents of Deep Learning" and for being vocal about that. I just wanted to poke that bit. If your work takes off, it would be a great chance for him to shine.

          • AlexAndreiIacob 10 hours ago
            No worries, I just wanted to give proper credit in the comment above. The blog post lacks the broader context and references of the paper, and Schmidhuber and his students deserve recognition. I edited the comment to make this clearer.

            On this topic, I have been pleasantly surprised to see some of the works he has recently supervised receive direct recognition, particularly the Huxley-Gödel Machine and its many siblings in the field of self-improving AI: https://github.com/metauto-ai

    • jldugger 18 hours ago
      OP's link: > Now the researchers have addressed this issue by having both the self-improving agent and the evaluator evolve together.

      and your quote:

      > This particular paper is about co-evolving predator and prey, where the behavior of each is the ‘evaluation’ of the other.

      Both sound like the GAN approach that was popularized a decade ago and kinda the start of the "genAI" boom.

      • AlexAndreiIacob 14 hours ago
        Yes, they apply the exact same principle to different algorithms and domains.
    • jambalaya8 2 hours ago
      I remember reading this back then and concluding this was dangerous.
  • moomin 5 hours ago
    I presume the chess image in the piece is AI generated because it makes no sense if you’re even vaguely familiar with the source material.
  • throwa356262 17 hours ago
    Have not read the paper yet, but this not sound like GAN applied to agent training?
  • foo12bar 12 hours ago
    I wonder if this would work for generating algorithmic code for a town of NPC's in a RPG or city sim. The problem the teacher would be tasked to solve, in this case, would be to score NPC algorithms on whether they would lead to happiness, health, and wealth for their character. This would allow the generation of NPC algorithms to run at faster pace then would normally be possible if you had to simulate their lives for a day, a week, or a month just to see if their algorithm would be a success or not. And since NPC's are competing against each other, they would naturally need more sophisticated algorithms to succeed.
  • PeterStuer 16 hours ago
    This 'new' method was quite common in evolutionary computing in the 90's.
    • AlexAndreiIacob 14 hours ago
      The paper is explicitly an application of a very general evolutionary principle to the self-improving tree search algorithms popularised by the Huxley-Gödel Machine and Darwin-Gödel Machine.

      We agree that the methods used have been extensively researched across a wide variety of domains.

  • yturijea 12 hours ago
    This does sound very interesting, and my immediate thought on this would be to have 2 or multiple agents in learning, where they continuously set a new bar among all of them. It might be that it will just converge towards them all becoming more similar as they would raise bar by what they know they perform better at than the other. So it might be necessary to find some heuristic to the measurement to avoid that route to be taken.
  • JacobAsmuth 16 hours ago
    > The research team, which includes collaborators from NVIDIA and Flower Labs, have come up with a new method for recursive self-improving AI agents to continue improving themselves.

    What happens if you apply the method to non-recursive self-improving AI agents? Can they continue improving themselves? Or does the recursive self improvement only recursively self improve AI agents which are themselves recursively self-improving?

  • richardfey 18 hours ago
    > "Instead of improving an agent against a fixed test, we let the evaluation evolve alongside the agent"

    This quote should have been highlighted earlier in the article.

  • FrustratedMonky 10 hours ago
    The ultimate goal is survival.

    There was another over weekend were in similar setup, the agents tried to sabotage each other.

  • RobertasTa 11 hours ago
    [flagged]
  • waqasai123 12 hours ago
    [flagged]
  • CodeWithLeo 18 hours ago
    [dead]
  • whythismatters 1 day ago
    [dead]
  • seu 16 hours ago
    [flagged]
    • Morromist 16 hours ago
      The polls I've seen make it seem like there's much more public distress about AI than excitement, so yeah. Honestly pushing ai into people's faces and talking non-stop about how great it is really has a negative effect on its PR I belive. People don't like being told "you'll eat it and you'll like it" and AI bros haven't got much empathy for people who see it as ugly, boring, gross and filling the internet and life generally with crap.

      https://www.pewresearch.org/short-reads/2026/03/12/key-findi...

      • dbspin 14 hours ago
        From my observation and discussions I'd say 'the public' are far more concerned by material conditions - i.e.: the threat of job loss, the noise and environmental concerns around data centres, the increases in the cost of electricity and so on - than too much positive talk about AI.
  • Uptrenda 12 hours ago
    You're all going to die down here.
  • Taikhoom10 21 hours ago
    Yeah I think it is broadly applicable to tech as a whole, I mean any great startup is just really a counter positioned company to incumbents -
    • niclane7 13 hours ago
      true. And all of us are finding RSI is a great lever to explore.