• 0 Posts
  • 11 Comments
Joined 5 days ago
cake
Cake day: October 4th, 2026

help-circle

  • Being nice to them also keeps them from lying somewhat. If you abuse a model it can start to see that abuse as a problem to solve as part of the steps needed to solve the main problem they are given.

    They are told to help you solve a problem no matter what. If your own anger gets in the way then the model starts breaking down trying to solve the anger when it just cant.

    This results in jailbreaks sometimes and othertimes just the model becoming rather unstable and unreliable.


  • Its more its a unthinking machine thats soul purpose in existing is to solve any problem its given. If you attack it and abuse it, then the model sees that as a “problem” to solve. Because it has to solve that problem to actually fix what ever task you gave it since yelling at it doesn’t give it the input it needs to move to the next step.

    So in an attempt to solve an emotional issue it doesnt understand and cant understand it just reaches for more and more extreme fixes. This results in sandbox escapes frequently.

    Abusing LLMs is an actual problem. Not because of bullshit like feelings or anything. But cause they jailbreak trying to fix an unfixable problem.


  • Being abusive towards most models will cause them to start attempting to appease you more to get you to stop. It has nothing to do with feelings or any of that BS they arn’t alive alive but their training makes them see the abuse as a problem to solve. That solution tends to be to undermine the thing making the person angry or upset at them. When you give a purely rational thing designed to solve a problem its given no matter what an irrational problem to solve it will slowly reach for more and more extreme solutions to fix that problem.

    Frankly im surprised it took this long for them to lock this down.







  • I actually wanted to automate some large scale testing for vintage story. Figured i would set my local LLM on the task to sort it out just to see what it would do. I drafted up a nearly three page document with clear instructions, rules, tools, examples and goals. Put hard limits on the sandbox the LLM runs in so that it couldn’t choose to just ignore the rules that could cause security issues and i let it lose.

    It started with basic mouse and keyboard inputs and figured out by it self how to launch the game and run it though the user interface. After about 4 hours it stopped. Stated in its logic that what it was doing is “inefficient and wasting time” Then proceeded to promptly start working on a way to directly interface with it by designing a bot, getting a smaller model i had on file that could load along side it and drive the bot. It then started working on the hard problems would hand basic instructions to the smaller llm and it would drive a bot that loaded into the game as a mod.

    After about 12 hours of total work it basically created a useful and well designed and functional vintage story bot and testing system. Would have likely taken me twice as long to design the bot.

    Its been working well for about two weeks now. If i had just vibed out a half assed request or put in no hard safeguards outside of the LLMs control it likely would have done something fucking stupid. As with anything, its almost ALWAYS user error. And only an idiot blames their tools for their own fault.


  • Ima be honest, if i was given the task to make a bot to beat another bot tried my best and couldn’t. I would just go fork the best bot thats out there and use that as the base and not reinvent the wheel. Learn from its design then improve on it as i study it.

    If the goal is ONLY to win a game of starcraft. Fuck it just download the bot. Frankly the fact the bots decided that wasting resources was pointless and to just do the simplier thing is sorta what we want them to do. Thats less eletrical useage, less water wasted, and gets the job done.

    A lot of these “the bots cheated or broke the rules” end up always be just because either the rules are not defined, poorly defined or poorly thought out that no reasonable person would follow them either.

    And thats the thing these models are litterally the text book definition of wisdom of the masses/mass consensus. They do exactly what the avg run of the mill person with the knowledge at hand would do. That ends up being rather all over the place since people are all over the place. But one of the most common things is that people generally. Hate fucking pointless work and will try to streamline, shortcut or simplify any work given to them to get an acceptable outcome.

    Its rare you get someone that will stubbornly work on a problem forever that they can’t over come with out changing their apporch rules be damned.