
I randomly came across this video on YouTube and then read a post where someone described PRAXIST as “10x Autoresearch ”
That felt like something worth testing. Autoresearch by Andrej Karpathy is easy to understand. You give an agent a model-training task, let it try changes, run the experiment, check the score, then try again. Praxist takes that idea further. It can run several research directions at once, keep a record of what worked and what did not, and use earlier results to decide what to test next. I wanted to see what happened when I gave it a normal machine-learning problem instead of asking it to optimize a toy demo.
The Plan
I used the Adult Income dataset, a standard tabular ML problem. The model has to predict whether a person falls above an income threshold using information such as education, job type, age, hours worked, and capital gains.
I kept the starting model simple: basic data cleanup, category encoding, and logistic regression. The agents were allowed to experiment with preprocessing and feature engineering. They were not allowed to change the scoring method, dataset, or evaluation rules.
What it found
One of the better ideas was also one of the simplest - create extra numerical features that let the model consider relationships between numbers instead of looking at each one separately. I ran it for a short period with a modest budget of 2M tokens. Here is what it found:

It created features such as:
- hours-per-week² — whether working 60 hours has a different effect than simply twice the effect of 30 hours
- age × education-num — whether education has a different relationship with income at different ages
- education-num × hours-per-week — whether education and working hours together matter more than either alone
- capital-gain × capital-loss — whether the combination of investment gains and losses carries useful signal
- age × capital-gain — whether investment income has different meaning across age groups
The degree-two version scored better than the baseline on each of the five configured development runs. That makes it a useful result to carry forward. Some preprocessing and a feature-selection approach also made things worse. One combination that sounded reasonable on paper performed badly enough that there was no reason to keep testing it.
That was the part I liked. The system did not only give me one better number. It gave me a record of which directions were useful, which ones were not, and which ones still needed more testing.
Why this is promising
A lot of AI-agent demos end at “the agent changed code and the score went up.” That is fine as a demo, but it is not how I would want to trust a long research project. If an AI system is going to run many experiments, I want to know answers to questions like
- What did it change?
- What was it trying to test?
- Did the result hold up when rerun?
- What failed? What should it avoid repeating?
- Can a person look at the final result and follow the path that led there?
Praxist is trying to make those questions part of the workflow instead of leaving them to someone reading logs afterward.
This is not a claim that Praxist has solved machine learning research. It is not a claim that the small improvement I found will work on every dataset. It is not a comparison against every AutoML tool or every other agent system.
It is simply a first proper test: a constrained ML task, a fixed way to score results, repeated runs, and a record of both the useful and useless experiments. As a first run, I think the basic idea held up. Praxist did more than generate a suggestion. It helped turn a set of experiments into something I can inspect, rerun, and build on.