TellMe and then ByeByeV2
ByeBye
ByeByeV2 is a GUI take on my most popular GitHub project ByeBye. ByeBye is a TUI application which lets you select from a set of system power functions. “What is the use case for that?”, you might ask. Imagine a scenario where you are working on a window manager Linux desktop that is based on your favorite Linux distro. After a long day of work, you want to get a coffee, but first turn off your computer, or, better yet, hibernate. You don’t know who you might meet at the coffee machine, and these coffee machine talks can be loooong. So you are still in the terminal and want to hibernate or log out or what have you … but wait! … what is the command again? Or what is the keymap again? “I usually just hit the power button,” you think. And then you start typing: “hiber<tab><tab>” … nothing happens.
This is the use case! I want to remember one command, one application that I can launch in the blink of a second and that is “byebye”. The reason I wanted to make ByeByeV2, which is a GUI (Graphical User Interface) rather than a TUI (Terminal User Interface) is to test how quickly I can reproduce the ByeBye TUI application as a platform native GUI application with current frontier AI models.
TellMe
While trying to analyze what the agent created, I found a problem: How can I look at the code and at the same time look at the agent’s explanation of the code without constantly switching between the agent window and the editor window. Furthermore, in addition to this physical problem, I had a logical problem: the agent will always give me a high level explanation of what is going on in the application. When I ask it how the Swift side of the application works, it’ll tell me that there are three files and the program starts with this first file and so on. Looking at the code, I then ask myself, “but how does this file do that?” Going back and getting a line-by-line answer is very annoying. So I thought, can I make a Vim plugin that shows me the explanation in a float? Then I thought, why not make it an application independent of Vim? So I asked an agent to write me an application which can explain another application to me with the help of another agent (fun times, fun times). The result was TellMe and the question that I asked myself was the same as with ByeByeV2. What did the agent do? Would I do it the same way? Is this the optimal way?
Constraints
To be fair, I gave the agents more than just the application requirements. For ByeBye, I gave it a screenshot of the TUI application and told it not to use any dependencies other than the platform APIs. For TellMe, I gave it a very specific set of requirements and told it not to use any dependencies except for Jinja2, which it started reimplementing in the process of building its application. I told it not to use dependencies because of security reasons. In my experience, once you leave the agent to do whatever it pleases, it goes crazy with lots and lots of dependencies. I want to vet those first, but ideally I don’t even want to bring those in. Especially for personal software I don’t want to download a lot of packages for an application that I might use once or twice.
Do these constraints affect the results of this experiment? Yes, they force the agent to implement all components of the system. In a very complicated application, this might cause problems because every component first has to be right before the entire system can work. In simple applications like ByeByeV2 and TellMe this does not impact the software system that much. Most things you need are installed anyway. That means Swift and SwiftUI for Mac applications, GTK for Linux, and the Win32 API for Windows. For TellMe, the agent chose to use Python and JavaScript/CSS/HTML and benefited from the broad standard library that Python brings to the table.
What did the agent do?
ByeBye
The agent built a nice architecture with a core application part in C and a UI part in Objective-C. Both are well separated from each other. What tripped me up with this application is that the agent opted for Objective-C rather than Swift, even though Swift with SwiftUI is the standard for Mac applications. A second prompt, however, translated the Objective-C application into Swift. No problem at all, even though I never wrote or read a line of code in either language.
The core application and the Linux GUI are written in C. The C code is reasonable, except for a slightly excessive use of ternary operators for my taste. One particularly good part of the GTK application is how it detects the Linux desktop’s capabilities and builds the GUI accordingly, so it doesn’t expose options that can’t be executed. That said, the application is very, very simple.
Other than that there is not much to say. I am surprisingly OK with the way the agent wrote ByeByeV2.
TellMe

This screenshot shows the TellMe application. It is stunning what is possible with two or three directed prompts. But this application, in contrast to ByeBye, was not smooth sailing.
Problem: To be super fancy, the agent used Mac sandbox-exec to run
OpenCode, which I used as a proxy to not have to handle auth and other annoying
setup. Simply put, OpenCode does not run correctly in a sandboxed environment on
Mac. This was not the real problem; the actual problem was that the agent tested
the application and claimed it would work. Hidden in the complexity of the test
and the implementation was the fact that the implementation was wrong. Which
This is surprising when you don’t know how the interface to OpenCode was implemented.
ChatGPT Astra used up my five-hour session token capacity, leaving me unable to
continue for 30 minutes. I would have spent the same amount of time doing it by
hand, but it would have worked. I don’t know what I should expect when starting
an agent on something, but I would expect that
calling a well-documented CLI would not be that hard. Then I had to spend
another 30 minutes fixing the problem. Though the rest of the application would
have taken me much more time to build by hand, would have looked worse, and would have had
fewer features, this particular part was too complex and I would have built it in
minutes on my own and it would have worked.
The difference between the 150 lines of code for the OpenCode CLI provider and
the rest of the agentic work is that the agent made some wrong assumptions about
the inner workings of sandbox-exec and OpenCode. This was the only part with
hidden complexity. The rest of the application is straightforward:
- CLI Arguments
- Web Server
- SQLite3
- File System Interaction
- Templating
- JavaScript/CSS/HTML (UI system)
Even though it is straightforward, the verbosity of the code is remarkable. It fulfills the requirements but it also extends the code here and there to catch some corner case or add some configuration that is not needed. These corner cases perceived by the agent (that actually aren’t), configurations, and considerations that a human developer would just leave out and add on a need-to-have basis rather than in advance lead to prop drilling and other types of added complexity that is hard to understand as a human with limited RAM. This additional complexity moves through the entire application. Sometimes this is so extensive that there is dead code in a completely new application with little complexity. There are still some problems even with simple applications that need to be fixed in agentic code. I can imagine that the problem is that the agent has too few constraints in a greenfield project. A project that has no infrastructure does not provide a clear solution path from the context of the project.
The cool thing about the application is that it took only about 3 hours to build, and it has a fair amount of functionality and it is actually useful for my personal use case. This is the thing with personal, one-off software: it does not have to be perfect, it does not have to look super good, it does not have to be maintainable, it can just be built fast, extended, used, and discarded. The current models are also capable of choosing a good interface and architecture for a simple application like this. I would question whether this is still the case for bigger applications with more inherent complexity, but these small applications are no problem anymore.