This page follows one real AIUS session end to end, run with the released terminal against the public Bank Marketing dataset. It took 12 minutes and $0.97 of model usage. The numbers, the report and the final summary come from that session. The prompt and the agent at work were retaken with a later release, entering the same prompt on the same data.

The question

A retail bank’s direct-marketing team sells term deposits by phone. About 11% of past contacts subscribed, so most calls are wasted. The team wants to call the most promising clients first. The project folder held three things: the data, its data dictionary, and a short README.
term-deposit-campaign/
├── README.md
├── data/
│   ├── bank-marketing-campaigns.csv   41,188 past contacts, 21 columns
│   └── data_dictionary.md
└── .venv/                              pandas, scikit-learn, duckdb, matplotlib
The data is Bank Marketing (bank-additional-full) by S. Moro, P. Cortez and P. Rita, from the UCI Machine Learning Repository, licensed CC BY 4.0. The prompt, typed into the terminal in that folder:
The prompt in the AIUS terminal: build a propensity model with honest cross-validation, compare it with a baseline, report ROC-AUC and top-10% reach, explain the drivers, save code under output/ and publish a report to a new project in AIUS Examples

What the agent did

1

Read the brief and checked the environment

It read README.md and the data dictionary, confirmed the project’s .venv had scikit-learn, pandas and DuckDB, looked up the dataset’s source page and created the project in AIUS Examples.
2

Profiled the data before modelling

It found 41,188 rows, no missing values, 12 duplicate rows and 11.27% of contacts subscribing. It flagged duration (the length of the last call) as leakage: it is only known after the call it is meant to predict. It also treated pdays = 999 as “never contacted before”, not as a number of days.
The terminal while the agent works: it has read the README and data dictionary, inspected the Python environment and profiled the dataset
3

Noticed that time matters

The rows are in the order the calls were made, and subscription rose from 2.8% of contacts in the first tenth to 45.9% in the last. A random train/test split would mix past and future and overstate performance, so the agent used forward validation: four folds of 6,590 contacts, each trained only on earlier rows, then one untouched final period of 8,238 contacts.
4

Kept only what is known before a call

The model uses the client profile and previous-campaign history. It leaves out the call’s duration, the current campaign’s contact count, the channel and calendar fields, which are not known when the call list is drawn up. Economic indicators were tested separately as a sensitivity check.
5

Compared against baselines and checked its own numbers

It compared logistic regression with gradient boosting and a random ordering, then wrote a separate verifier. The verifier recalculated all 15 ROC-AUC and top-decile results independently and reproduced the full run from scratch with identical results.
6

Asked before publishing

It proposed attaching the two scripts and three result files, and keeping the raw data, row-level predictions and model file on the computer. After approval it ran the report review (no warnings), published the report, read it back and gave the link.

The result

Forward CV is the mean ROC-AUC, ± its standard deviation, over the four forward folds. Final period is the ROC-AUC on the untouched last 8,238 contacts. Top 10% reach is the share of that period’s subscribers found among its top-scored 10% of contacts.
ModelForward CVFinal periodTop 10% reach
Random ordering0.5000.50010.0%
Logistic regression (selected)0.531 ± 0.0490.68522.5% (571 of 2,540)
Gradient boosting0.518 ± 0.0490.67618.9%
Calling the top-scored 10% of the final period’s contacts would have reached 2.25 times as many subscribers as calling at random. Previous-campaign history was by far the strongest signal, followed by job. The agent’s recommendation was a monitored pilot, not a rollout: early folds were close to random, and the population changed a lot over time.
Published scores for this dataset vary a lot with the validation design and the features allowed. This analysis kept to data known before a call and validated forward in time, so its scores are lower and more cautious; the report explains why.

The published report

The report leads with the decision and its uncertainty, then shows the data’s limits, the model comparison, the drivers and the method. Its numbers carry footnotes to the attached files they came from.
The report headline, recommending logistic regression for a monitored pilot, with its summary and five deliverables
The Data sources card: 41,188 rows, 16 of 21 fields used, 4 alerts, and a statement that the data cannot establish current-client performance or incremental subscriptions
The Model comparison section: a table of cross-validated and holdout ROC-AUC and top-10% recall for random ordering, logistic regression, gradient boosting and a boosting sensitivity check
A chart of the share of final-period subscribers reached against the share of contacts called, for logistic regression, gradient boosting and random order
The attached files are in the project’s Files panel and on its Outputs page, and the project’s Work view records the published report. The screenshots on those pages are of this project. Anyone in the organisation can download the code and rerun it:
.venv/bin/python output/propensity.py --out output/new-run
.venv/bin/python output/verify.py output/new-run

Try it yourself

Download the dataset from the UCI page above, put it in data/ with a short README, create a .venv as in Installation, and give the agent the same prompt with your own organisation’s name. Expect different wording; the agent works through the same checks.