• 0 Posts
  • 8 Comments
Joined 3 years ago
cake
Cake day: June 19th, 2023

help-circle

  • You get the LLM to write code quickly and then you review it and perform experiments manually.

    I think this is where the issue is, the word “review”.

    Reviewing comes in two different types: non-trivial reviews and trivial reviews.


    Non-trivial change example: Create a presentation where the user flow follows a flow chart.

    Someone could go to an LLM, prompt it with “create a presentation that follows this flow chart” followed by the mermaid syntax of the flowchart.

    The LLM will give you back an array of slides where certain functions/actions/triggers/etc… navigates to a different slide based on its index within the array.

    But if you ask a person to do it, they might sit there for a while, try a few different attempts to understand the problem better, then come back to you with some typed generics and a map/dictionary/object/associative-array/etc… with slides in it, and functions/actions/triggers/etc… that navigate to a flow chart by a slide’s id.

    Two different bits of code to review from two different sources.

    You can choose to do one of the following:

    • “LGTM” the changes (in which case it wasn’t actually reviewed),
    • Read through the entire change to try and comprehend it until you find a part that you don’t understand
      • Sidenote: If you didn’t find a part that you didn’t understand, then the change isn’t actually non-trivial, so you can refer to the “trivial change example” below. This section is about non-trivial changes.

    When asking the LLM a question about the part that you didn’t understand, it will either give you:

    • A completely different changeset, so now you have something completely different to review. And by the fact of the longest part of programming being digesting code you didn’t write, in effect you’ve taken a long-cut and could have written it yourself faster. (This is part of where the “LLMs make programmers take longer” observation comes from)
    • A post-hoc justification for it. Which would by it’s nature not have take place before the code was written, which makes it unable to have actually affected the code, and thus not actually be a valid reason for why the code is the way it is. So you don’t get a valid answer, you just get a convincing one.

    When asking the person why they did it, they’ll tell you they tried a few attempts to get their head around it, and mid-attempt they accidentally commented-out one of the slides which created an unseen error when one of the functions/actions/triggers/etc… tried to go to a slide that didn’t exist. So to prevent that problem from occurring again, they wrote another attempt where they used types/generics such that every slide’s id and reference to every slide’s id was type checked. That way, if a reference was ever incorrect or initially correct but made incorrect by a later change somewhere else, the editor would alert you before you even tried to compile your code. With that response, you now have the reasoning behind the non-trivial thing you didn’t understand.


    Trivial change example: Changing a color from “orange” to “red”.

    Someone could go to an LLM’s chat window and type, “Change the color to red”.

    But it’d be faster to just double click the word “orange” to select it, then type the word “red”.

    So for trivial examples, it doesn’t really make much sense to use an LLM, it’s literally faster to do it yourself then review your own trivial changes in a diff.


    Coming back to the problem of the word “review”:

    You get the LLM to write code quickly and then you review it and perform experiments manually.

    • Non-trivial reviews require back and forth communication. Reviews can be convincing without being valid. The review’s validity depends upon the validity of reasoning within that communication.
    • Trivial reviews are trivial, so there’s no point to using an LLM in the first place

    That means that reviews of LLM outputted code by their nature are either invalid and/or non-optimal.

    My advice: cut out the crutch/middle-man and do the hard work of establishing that rock-solid understanding. You’ll be much better off in the long run.


  • But that’s what I already have and it doesn’t work properly.

    Here’s my minimal test case for a typical mid-work layout:

    • Open a window
    • Open 4 new tabs
    • Type the tab’s number in each tab so it’s easy to tell them apart when trying to switch between them
    • Drag tabs 3 and 4 to the right side of the window so that the result is a window with 2 panes each of which contain 2 unique tabs.
    • Click on Tab 1 to select it

    Now the test:

    • Press Ctrl + Tab to switch to tab 2. Works
    • Press Ctrl + Tab to switch to tab 3. Does not work, tab 1 is selected instead

    The tab switching is stuck switching only the tabs in that pane, counter to the way every other tabbed program works.

    Holding Ctrl + Tab should cycle through every tab.


  • So just because I didn’t write a book means I could never understand it from reading it?

    Reading a book does not give you the same knowledge as the author.

    The author may have had to choose between two different pieces of exclusive information to add to the book, they may have written information then had to remove it for some reason, or may have had to exclude certain information it altogether.

    You would never know if they did, or the reasons behind why they did what they did if you only read the book.


    Are editors useless in the world of book publishing? Then what is the point of code review?

    The difference between your book analogy and human code review is you can actually talk with the author in a bi-directional channel to gain a deeper understanding of what they did, and why they did it.

    Book editors and human code reviewers are the same in that respect, they develop the deeper understanding that simply reading does not provide.


    When it comes to LLMs: that bi-directional channel doesn’t exist, only the illusion of one.

    LLMs don’t have the capacity to think things through, their reasoning is actually post-hoc justifications for their previous output, not a result of a priori thought. There is no actual understanding that could be gleaned because it literally doesn’t exist.

    To train an LLM you need more data and feedback than can be manually tagged, so tags and feedback are generated automatically. This means that they are not being trained against ground truths, they’re being trained against a confidence checker. To an LLM, there is literally no difference between a correct answer and a confident answer.

    That’s why they seem so stupid when they give an answer that is obviously incorrect.

    They are confidence machines, they produce confident sounding answers.

    The problem is that human psychology is wired not to be discern the difference when the output’s falseness isn’t immediately obvious.


  • I tried to switch to Zed at the start of the year but couldn’t Zed’s tabs to work.

    It seems like the biggest paper cut preventing users from switching to Zed and they don’t seemed interested in fixing it.

    I couldn’t find any settings to fix it.

    Were you able to figure it out?

    How every program (e.g. editors, browsers, file explorers, etc…) works:

    • Ctrl + Tab switches to the next tab
    • Ctrl + Shift + Tab switches to the previous tab

    How Zed works:

    • Ctrl + Tab and Ctrl + Shift + Tab brings up a tab switcher widget which doesn’t even list all tabs

  • Not always.

    Yes always. To paraphrase Feynmann: What you did not create, you have not understood.

    Prototyping can be valuable in many circumstances when you don’t know the shape you’d like the code to take.

    Yes

    Sometimes you need to try something before you know if it’s right.

    Yes

    I’d rather throw away prototype code written by an LLM than code I had to write the hard way.

    No.

    If you consider writing your own code “the hard way”, then your goal should be becoming better at prototyping yourself. It’s a skill that you need to requires purposeful effort to strengthen. But once you have it, you get both the understanding behind the code and the prototype.