AI Agent Capabilities

Why an AI agent in finnts?

The AI agent is a tool-calling orchestration layer that sits on top of the core finnts pipeline. It uses an LLM to:

You keep control through a few inputs (data, horizon, optional regressors, performance goal, iteration budget); the agent does the rest.

Core Agent Functions:


What the agent produces

Use these helpers to retrieve outputs:


Prerequisites

Set up environment variables for Azure OpenAI (example):

Sys.setenv(
  AZURE_OPENAI_ENDPOINT    = "<your-endpoint>",
  AZURE_OPENAI_API_KEY     = "<your-key>",
  AZURE_OPENAI_API_VERSION = "<api-version>"
)

End-to-end: first run with the AI agent

Below is a complete flow using the built-in M4 monthly sample.

1) Create a project

library(finnts)
library(dplyr)

project <- set_project_info(
  project_name = "ai_agent_demo",
  path = tempdir(), # or a persistent folder
  combo_variables = c("id"),
  target_variable = "value",
  date_type = "month", # day|week|month|quarter|year
  fiscal_year_start = 1 # fiscal month (1 = Jan)
)

Tip: path controls where logs/forecasts/EDA artifacts are saved.
Supports local filesystem, Azure Blob (via AzureStor::blob_container), or Microsoft 365 drives (ms365r) via storage_object.

2) Bring data

hist_data <- timetk::m4_monthly %>%
  dplyr::filter(date >= as.Date("2013-01-01")) %>%
  dplyr::rename(Date = date) %>%
  dplyr::mutate(id = as.character(id))

3) Define the LLM

llm is an ellmer Chat that Finn uses as a configuration template. Finn never mutates this template. It creates an isolated, empty-history session for each forecast series and each ask_agent() request, so one modern model can perform both forecast input selection and results analysis without conversation history leaking between workflows.

llm <- ellmer::chat_azure_openai(model = "gpt-4o-mini")

4) Create the agent run

agent <- set_agent_info(
  project_info = project,
  llm = llm,
  input_data = hist_data,
  forecast_horizon = 6, # number of future periods
  external_regressors = NULL, # e.g., c("Price","Promo")
  allow_hierarchical_forecast = FALSE, # set TRUE to let agent use hierarchies
  negative_forecast = FALSE, # set TRUE to allow forecasts below zero
  overwrite = TRUE # start a fresh run_id if inputs changed
)

This writes the versioned inputs into path/input_data/ (hashed by combo/run) and logs the new agent_version/run_id.

5) Let the agent iterate to a best run

iterate_forecast(
  agent_info          = agent,
  weighted_mape_goal  = 0.05, # your accuracy target of 5%
  max_iter            = 3, # stop after N iterations if not hitting goal
)

What happens under the hood:

6) Retrieve results

best_runs <- get_best_agent_run(agent_info = agent, full_run_info = TRUE)
head(best_runs)

fcst <- get_agent_forecast(agent_info = agent)
head(fcst)

Ask questions about your forecast results

After running iterate_forecast() or update_forecast(), you can use ask_agent() to ask natural language questions about your results. The agent analyzes your forecast data, model configurations, and EDA outputs to provide data-driven answers.

How it works

ask_agent() creates an LLM-driven workflow that: 1. Plans the analysis steps needed to answer your question 2. Executes R code to analyze the relevant data 3. Generates a natural language answer based on the results

Example questions

# Ask about forecast accuracy
answer <- ask_agent(
  agent_info = agent,
  question = "What is the average weighted MAPE across all time series?"
)

# Ask about models used
answer <- ask_agent(
  agent_info = agent,
  question = "Which models were selected as best for each time series?"
)

# Ask about feature importance
answer <- ask_agent(
  agent_info = agent,
  question = "What are the top 3 most important features for the forecast models?"
)

# Ask about data quality
answer <- ask_agent(
  agent_info = agent,
  question = "Were there any missing values or outliers in the data?"
)

# Ask about specific forecasts
answer <- ask_agent(
  agent_info = agent,
  question = "What are the forecasted values for M750 for the next 3 months?"
)

# Ask about time series characteristics
answer <- ask_agent(
  agent_info = agent,
  question = "Which time series show strong seasonality patterns?"
)

# Ask comparative questions
answer <- ask_agent(
  agent_info = agent,
  question = "Which time series have the highest forecast uncertainty?"
)

What data sources are available

ask_agent() has access to four main data sources:

  1. Forecast results (get_agent_forecast()): Future predictions, back-test results, model selections, confidence intervals
  2. Model configurations (get_best_agent_run()): Feature engineering settings, transformations applied, model hyperparameters
  3. EDA results (get_eda_data()): Time series characteristics, seasonality, stationarity tests, data quality metrics
  4. Model summaries (get_summarized_models()): Feature importance, model parameters, recipe details

The agent automatically determines which data sources to use based on your question.

Tips for effective questions


Updating with new data (production loop)

When you have new input data, keep the same project and create a new agent run with updated input_data. Then call update_forecast():

# suppose you've appended more months to hist_data:
hist_data2 <- hist_data %>% dplyr::filter(Date <= as.Date("2016-06-01"))

agent2 <- set_agent_info(
  project_info = project,
  llm = llm,
  input_data = hist_data2,
  forecast_horizon = 6,
  overwrite = TRUE # required to create a new agent version when running update_forecast()
)

update_forecast(
  agent_info             = agent2,
  weighted_mape_goal     = 0.05,
  allow_iterate_forecast = TRUE, # if degradation detected, allow the agent to re-iterate
  max_iter               = 2 # cap re-iteration cost
)

updated_fcst <- get_agent_forecast(agent2)

# Ask questions about the updated forecast
answer <- ask_agent(
  agent_info = agent2,
  question = "Summarize the forecast accuracy."
)

What update_forecast() does:


Hierarchies (optional)

Set allow_hierarchical_forecast = TRUE in set_agent_info() to let the agent detect:

When the agent selects a hierarchy, it will: - train at the selected aggregate(s), - reconcile down to the bottom level, - produce a reconciled get_agent_forecast() output.

For background and manual control, see the “Hierarchical Forecasting” vignette.


External regressors (xregs)

If you pass external_regressors = c("Price","Promo", ...):


Parallelism knobs

Every time-series combo receives independent driver and reasoning Chat objects with empty conversation history, whether execution is sequential or parallel. Parallel runs require ellmer 0.4.0 or later on the driver and every worker; Finn serializes the configured Chats through foreach before creating the per-combo deep clones.


Reading artifacts directly (optional)

You normally won’t need this, but for audits:

Use the helpers first; dig into files only if you must.