Parallel multi-tenant scoring

This tutorial shows how to run batch scoring for multiple independent tenants in parallel using temporary Dataiku Application instances.

In a multi-tenant setup, each tenant can store its source data in a separate BigQuery dataset while sharing the same table structure. For example, each dataset can contain tables for students, behavior, and attendance. You can use a single scoring flow and a global saved model to prepare this data and write predictions back to the corresponding tenant dataset.

A shared Dataiku project is not suitable for concurrent scoring requests because its variables and intermediate datasets are shared between runs. Dataiku Applications allow you to use the project as a template and create an isolated instance for each request. Each temporary instance has its own project variables, scenario run, and intermediate datasets.

In this tutorial, you will create a tenant-aware scoring flow, convert it into a Dataiku Application, and use the Dataiku API to create and run a temporary instance for each tenant. The scoring service sets the target BigQuery dataset on the instance, runs the scoring scenario, and deletes the instance once the run completes.

Prerequisites

  • Dataiku >= 12.0

  • Python >= 3.12

  • A BigQuery connection.

  • A global saved model that can score the feature table built by the flow.

  • At least two BigQuery datasets, for example, tenant_001 and tenant_002, each containing students, behavior, attendance, and predictions tables with the same schema.

dataset samples

If needed, you can download a small subset of these input datasets:

Tip

To create this dataset, go to the Google BigQuery console, create the dataset, and within it, create the tables.

Creating the tenant-aware scoring flow

Create a project named, for example, MTS_TEMPLATE. Build the flow as you would for one tenant:

  • Read the three source tables.

  • Prepare the data.

  • Create a feature table.

  • Batch-score it with the global model.

  • write the predictions to BigQuery.

Create a project variable named tenant_bq_dataset. For each input dataset and the prediction output dataset, open Settings, select Advanced, and scroll to Overrides. Under Override variables, add an override on params.schema with the value ${tenant_bq_dataset}. In BigQuery, the field called dataset is a schema-like container that links between the Google Cloud project to the table. For all other datasets, if they belong to the same tenant, you have to follow the same process. If they belong to another connection, make sure the datasets are relocatable (see Making relocatable managed datasets).

Set tenant_001 as the default value in the project’s Variables. You can now build the flow manually to validate it. For example, the Dataiku datasets then point to prod.tenant_001.students, prod.tenant_001.behaviour, and prod.tenant_001.attendance. The API caller will replace that value in the temporary instance when it is created.

Attention

Do not set tenant_bq_dataset in a shared project from a scenario. The variable is isolated only after the project has been instantiated.

Tip

If you need help in creating the project, you can follow this checklist:

  1. Create tenant_001 and tenant_002 datasets in BigQuery.

  2. In each one, create the identical tables: students, behavior, attendance, and predictions.

  3. Create the Dataiku project MTS_TEMPLATE.

  4. Add a BigQuery connection with access to all tenant datasets.

  5. Add the external BigQuery datasets students, behavior, attendance, and predictions to Dataiku; these are initially configured for tenant_001.

  6. Create the project variable tenant_bq_dataset=tenant_001.

  7. Add the override params.schema=${tenant_bq_dataset} to each of the four datasets.

  8. Build the Flow: preparation of the three sources -> feature table -> scoring using the global model -> predictions.

  9. Create the SCORE_TENANT scenario with a forced build of predictions.

  10. Manually test the Flow for tenant_001, then tenant_002.

  11. Convert the project to a Dataiku Application.

Creating the Application template and scenario

Create a scenario named SCORE_TENANT with a single Build step targeting the final predictions dataset. Enable force-rebuild so that each request reads the tenant’s current source tables and rebuilds the full pipeline.

From the project menu, select Application Designer, then click Convert into a visual application. Leave the parameters at their default values because the application will be instantiated through the API.

Scoring one tenant

For each tenant, the API caller creates a temporary Application instance, sets the instance’s BigQuery dataset variable, and runs the scenario. create_temporary_instance() returns a context manager, so Dataiku deletes the instance when the scenario run finishes, including when the run raises an error.

Code 1: Scoring tenants in parallel with temporary Application instances
from concurrent.futures import ThreadPoolExecutor
import dataikuapi

HOST = "https://<DSS_HOST>"
API_KEY = "<API_KEY>"
APP_ID = "PROJECT_MTS_TEMPLATE"
SCENARIO_ID = "SCORE_TENANT"


def score_tenant(tenant_id):
    client = dataikuapi.DSSClient(HOST, API_KEY)
    app_template = client.get_app(APP_ID)

    with app_template.create_temporary_instance() as app_instance:
        project = app_instance.get_as_project()
        project.update_variables({"tenant_bq_dataset": tenant_id})

        scenario = project.get_scenario(SCENARIO_ID)
        scenario_run = scenario.run_and_wait()
        if scenario_run.outcome != "SUCCESS":
            raise RuntimeError(
                f"Scoring failed for {tenant_id}, run {scenario_run.id}"
            )


tenant_ids = ["tenant_001", "tenant_002"]
with ThreadPoolExecutor(max_workers=len(tenant_ids)) as executor:
    list(executor.map(score_tenant, tenant_ids))

Note

To explicily choose the project key, use create_instance(instance_key, instance_name) instead. This creates a non-temporary Application instance. Delete its corresponding project with project.delete() when it is no longer needed. It may be useful to choose the project key if your underlying dataset storage has a length constraint on dataset names.

The prediction dataset uses external BigQuery storage, so it remains available after the temporary project is deleted. Ensure that all external intermediate datasets use locations unique to the tenant or instance before running the Flow in parallel.

Controlling concurrency

The example uses a ThreadPoolExecutor to demonstrate simultaneous requests. In a production API service, set max_workers to a value that matches the instance’s capacity and expected workload.

Attention

Temporary instances isolate project variables, but they do not remove platform limits. Concurrent work is still bounded by available jobs, activities, connections, and computing resources. Load-test with representative tenant data before increasing the caller’s concurrency limit.

Wrapping up

Congratulations! You can now score independent tenants in parallel without duplicating the flow definition or sharing project variables between runs.

Among the possible next steps, you can:

  • Persist a request ID, tenant ID, timing, and outcome in an external table.

  • Add an idempotency key to prevent duplicate scoring requests.

  • Return the prediction table location and run outcome to the calling application.

Reference documentation

Classes

dataikuapi.dss.app.DSSApp(client, app_id)

A handle to interact with an application on the DSS instance.

dataikuapi.dss.app.DSSAppInstance(client, ...)

Handle on an instance of an app.

dataikuapi.DSSClient(host[, api_key, ...])

Entry point for the DSS API client

dataikuapi.dss.project.DSSProject(client, ...)

A handle to interact with a project on the DSS instance.

dataikuapi.dss.scenario.DSSScenario(client, ...)

A handle to interact with a scenario on the DSS instance.

dataikuapi.dss.scenario.DSSScenarioRun(...)

A handle containing basic info about a past run of a scenario.

dataikuapi.dss.app.TemporaryDSSAppInstance(...)

Variant of DSSAppInstance that can be used as a Python context.

Functions

create_instance(instance_key, instance_name)

Create a new instance of this application.

create_temporary_instance()

Create a new temporary instance of this application.

delete([clear_managed_datasets, ...])

Delete the project

get_app(app_id)

Get a handle to interact with a specific app.

get_as_project()

Get a handle on the project corresponding to this application instance.

get_scenario(scenario_id)

Get a handle to interact with a specific scenario

run_and_wait([params, no_fail])

Request a run of the scenario and wait the end of the run to complete.

update_variables(variables[, type])

Updates a set of variables for this project