A roadmap I can build in public

Moving from analytics toward AI feels more useful when learning produces something I can explain and test. This is a proposed twelve-week personal roadmap for Soumyar, rather than a claim that I have completed these projects. The goal is to turn curiosity into small, inspectable experiments using public or synthetic data.

The first small experiment

The first project for this roadmap compares a linear forecasting model with two simple baselines on synthetic daily demand. It uses a chronological training/holdout split and records the prediction errors. The point is to practice reproducibility and evaluation before making claims about real-world performance.

Explore the forecasting experiment

Weeks 1–2: strengthen the foundations

Start with Python, pandas, basic probability and reproducible notebooks. Choose one public dataset, document what a row means, check missing values and duplicates, and produce a simple baseline analysis. The deliverable is a notebook that another person can run, with clear assumptions and a short explanation of its limitations.

Weeks 3–4: build one machine-learning baseline

Choose a small prediction problem and begin with a simple model. Separate training and evaluation data before fitting transformations; for time-dependent data, respect the order of events. Compare against a basic baseline, choose metrics that match the problem, and examine errors rather than reporting only a headline score. Google’s Machine Learning Crash Course is a useful foundation resource.

Weeks 5–6: explore language models

Learn how tokens, context limits and model instructions shape an application. Build a small summarization experiment using public text. Keep a fixed set of examples, compare outputs against the source, and record unsupported claims, omissions, latency and cost. Hugging Face’s LLM Course can support this stage after the Python and introductory deep-learning foundations.

Weeks 7–8: make a source-grounded assistant

Use a small collection of public documents to explore retrieval-augmented generation. Retrieve relevant passages and ask the model to answer with sources. Include questions that the documents cannot answer, and test whether the assistant acknowledges missing evidence. Retrieved text should be treated as data rather than trusted instructions. Keep private files and credentials out of the experiment.

Weeks 9–10: connect AI to analytics carefully

Build a read-only prototype that helps explain a synthetic dataset or proposes SQL for review. Use a small allowlisted schema and validate queries before execution; avoid giving the model unrestricted database access. Check whether explanations match actual query results. The useful outcome is a transparent workflow with known failure cases, not a chatbot that merely sounds confident.

Weeks 11–12: share the evidence

Turn the strongest experiment into a personal project page. Include the question, dataset source, approach, baseline, evaluation examples, limitations and reproducibility instructions. Remove secrets and unnecessary personal data before sharing. Describe what worked and what still needs improvement without presenting a prototype as a production system.

A pace I can adjust

The twelve weeks are a planning structure, not a guarantee. If the foundations need more time, extend them. A useful weekly rhythm is to learn one concept, build one small artifact and write one honest note about the result. Future articles can cover the notebook, the first model and the source-grounded assistant once there is real evidence to discuss.

Learning resources

Google Machine Learning Crash Course: https://developers.google.com/machine-learning/crash-course Hugging Face LLM Course: https://huggingface.co/learn/llm-course/chapter1/1