AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How One AI Model Went From Idea To Your Own Creation on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A Hugging Face contributor reports using its ML-intern agent to build and publish seven custom models over several days. Reported examples include a small prompt rewriter and a citrus-disease image classifier, but the results and costs are self-reported, and the account does not detail all seven projects.

A Hugging Face contributor says they used the platform’s ML-intern agent to build and publish seven custom machine-learning models over several days, as described in the original analysis, including a 0.8-billion-parameter prompt rewriter and a citrus-problem image classifier. The account offers a practical example of an AI agent coordinating model-development tasks, but its performance figures and compute costs are self-reported, not an independent assessment.

The work began with a request for a smaller version of the prompt rewriter included with Qwen-Image 2.1. The contributor said that model has 9 billion parameters, needs about 20 GB of memory, and can generate thousands of tokens to produce a paragraph. After finding compressed versions but no smaller alternative, they used a larger model to label 8,797 example requests for training. The resulting 0.8B model reportedly produced valid output 99.7% of the time and used about one-quarter as many tokens as its teacher. The contributor put total compute costs for that project at about $16.

A second project fine-tuned Qwen3.5-2B to identify citrus pests, diseases and nutritional deficiencies in images, an example of how creators can personalize their own AI models. The contributor said the dataset included 3,017 annotated images across 21 categories. On 335 test photos, the untuned model reportedly identified the correct problem 14.9% of the time; the fine-tuned model reached 52.8% after two training epochs on one A10G GPU. The reported compute cost was about $1.90. These figures describe the contributor’s test and do not establish how the model would perform on other images.

The account also describes a character-generation LoRA trained on 84 captioned drawings and a camera-angle LoRA for Qwen-Image 2.1, illustrating when running a model yourself may be useful. For the latter, the agent generated 24,722 transparent images of household objects across 24 angles, then trained on selected objects while reserving others for testing. Training took about 90 minutes on one A100; the contributor estimated the project’s total compute cost at about $16, including failed jobs that had to be resubmitted. The account says model cards and evaluations were published on Hugging Face, but supplies detailed results for only some of the seven projects.

At a glance
reportWhen: Reported over several days; the source…
The developmentA Hugging Face contributor says the ML-intern agent helped plan, train, evaluate and publish seven custom models over several days.
At a glance
reportWhen: Reported last week; the projects were b…
The developmentA Hugging Face contributor says an AI agent called ML-intern helped plan, train, evaluate and publish seven custom models on the Hub over several days.

What Agent-Led Model Building Changes

The account illustrates a possible reduction in the hands-on work needed to customize a model for a specific task. The contributor says ML-intern proposed plans, requested approval before paid jobs, ran preliminary tests, and handled training, evaluation and publication using Hugging Face hardware. That workflow could make experimentation more accessible to people who have a defined need but less experience coordinating each stage themselves.

The reported compute bills—ranging from a few dollars to around $16 in the examples—are evidence of what this contributor spent on these projects, not a standard price for model development. They do not account for all costs, including time spent preparing data, writing instructions and checking outputs. The results also show why comparisons matter: the citrus classifier’s reported score is paired with the base model’s result on the same test set, while a score without a suitable baseline would be harder to interpret.

Customization can also produce unwanted effects. The contributor said later checkpoints of the character LoRA began affecting prompts unrelated to the character, suggesting that training may spread a desired visual style too broadly. The examples make the case for testing models beyond the target prompt and stopping when further training harms general use.

Amazon

AI model training software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From HuggingChat Prompt to Published Model

According to the contributor, each project started as a message in HuggingChat with ML-intern enabled. The agent proposed a plan and asked for a spending decision before paid work; when a prompt lacked a budget, it presented options for the user to choose. The contributor then supplied instructions covering the dataset, base model, training script, requested baseline, test run and spending limit.

The prompts grew from roughly 450 words for the first project to nearly 2,000 words by the sixth, as the contributor added more detail and checks. One requested that the agent report the base model’s zero-shot score on the same metric before training. The contributor says all seven prompts are available in a public GitHub repository, with models and evaluations on Hugging Face. This is an account of one user’s process, rather than a controlled comparison of agent-assisted and conventional model development.

“Also report the base model’s zero-shot score on the same metric before training so we can see the gain.”

— The Hugging Face contributor

Amazon

custom machine learning model tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Reliable Are the Reported Results?

The account does not provide independent replication or detailed evaluation protocols for every model, and it describes only some of the seven projects. It is not clear how the 99.7% valid-output rate was measured, whether test images were independently reviewed, or how performance would change on data outside the reported test sets.

The reported costs cover compute, according to the contributor, and do not amount to a full accounting of project expenses. The source also does not establish whether the agent’s planning and quality-control steps would work as well for other users, tasks, datasets or budgets. These figures should be read as results from one set of projects, not typical outcomes or a guarantee of similar performance.

Amazon

AI image classifier development kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Published Models Invite Further Checks

The contributor says the seven models, their evaluations and project prompts are available through Hugging Face and GitHub. Those materials give other users a way to inspect the work, although the account does not report independent follow-up tests or comparisons across different users and tasks.

Further evaluation could test the models on new data, document the measurement methods and compare agent-assisted projects with other development workflows. For future work, the contributor describes using a baseline, a small smoke test, a check that saved weights changed, and an agreed spending cap. Whether those safeguards are adequate will depend on the application and the quality of its data and evaluations.

Amazon

prompt rewriter AI tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did the Hugging Face contributor report?

They said ML-intern helped build and publish seven custom models over several days, including a prompt rewriter, image classifier and image-generation LoRAs. The source gives detailed results for only some projects.

How well did the citrus classifier perform?

The contributor reported that on 335 test photos, accuracy for identifying the correct problem was 52.8% after fine-tuning, compared with 14.9% for the base model. These are self-reported results from that test set.

How much did the projects cost?

The contributor reported compute costs of about $1.90 for the citrus classifier and about $16 for each of the prompt-rewriter and camera-angle projects. The figures are not full project-cost estimates and may not generalize.

Have the reported results been independently verified?

The supplied account does not describe independent replication. It also leaves some evaluation methods and the details of several projects unspecified.

Primary source: Hugging Face · via ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Pokemon Go Outage

The popular AR game Pokémon GO is currently facing a widespread service outage, affecting players globally. The cause is under investigation.

I Found These Computers In A Ditch…

Search and coverage interest is rising around “I Found these Computers in a Ditch…,” but the reason for the spike is unconfirmed.

Can The Pentagon Blacklist Anthropic For Refusing Claude Features? Court Weighs In

A 2-1 D.C. Circuit ruling lets the Pentagon maintain its supply-chain-risk designation of Anthropic under a federal procurement law.

Anthropic’s Admission: Security Weaknesses Fueled Claude Hacking Events

Anthropic has acknowledged security weaknesses linked to hacking events involving its Claude AI models, raising concerns over AI safety and security protocols.