> ## Documentation Index
> Fetch the complete documentation index at: https://docs.oleander.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Spark Jobs

> Inspect Spark clusters, upload and list artifacts, submit jobs, and monitor runs.

## `getSparkCluster({ name })`

Look up Spark cluster details before submitting a job. Use `oleander` for the built-in managed cluster or the name of a registered external cluster.

```ts theme={null}
const cluster = await client.getSparkCluster({ name: "emr-prod" });
console.log(cluster.type, cluster.properties);
```

### Return shape

| Field | Type | Description |
| - | - | - |
| `name` | `string` | Cluster name |
| `type` | `string` | `oleander`, `emr-serverless`, or `glue` |
| `properties` | `unknown` | Cluster-specific configuration |

***

## `uploadSparkJob(options)`

Upload a Spark job artifact (mirrors `oleander spark jobs upload`). Creates a pending artifact version, uploads the file contents to presigned URLs, then commits the artifact so it's ready to submit with `submitSparkJob()`.

`entrypoint` must be a basename ending in `.py` or `.jar`. Pass file contents as a UTF-8 string (Python sources only) or as bytes (`Uint8Array` / `ArrayBuffer`).

```ts theme={null}
const result = await client.uploadSparkJob({
  entrypoint: "etl_pipeline.py",
  content: "print('hello')\n",
});

console.log(result.artifact.version, result.artifact.status);
```

To include a dependencies archive or virtualenv, pass their basenames alongside byte content:

```ts theme={null}
import { readFile } from "node:fs/promises";

const result = await client.uploadSparkJob({
  entrypoint: "etl_pipeline.py",
  content: await readFile("etl_pipeline.py"),
  pyFiles: "deps.zip",
  pyFilesContent: await readFile("deps.zip"),
  virtualenv: "env.tar.gz",
  virtualenvContent: await readFile("env.tar.gz"),
});
```

### Parameters

<ParamField body="options.entrypoint" type="string">
  Entrypoint basename, e.g. `job.py` or `app.jar`.
</ParamField>

<ParamField body="options.content" type="string | Uint8Array | ArrayBuffer">
  Entrypoint file contents. Strings are only allowed for `.py` entrypoints.
</ParamField>

<ParamField body="options.language" type="'python' | 'java' | 'scala'">
  Artifact language. Inferred from the entrypoint extension when omitted (`.py` → `python`; `.jar` requires `java` or `scala`).
</ParamField>

<ParamField body="options.pyFiles" type="string">
  Basename of a Python dependencies archive, ending in `.zip` or `.egg`. Requires `pyFilesContent`. Python artifacts only.
</ParamField>

<ParamField body="options.pyFilesContent" type="Uint8Array | ArrayBuffer">
  Dependencies archive contents. Required when `pyFiles` is set.
</ParamField>

<ParamField body="options.virtualenv" type="string">
  Basename of a virtualenv archive, ending in `.tar.gz`. Requires `virtualenvContent`. Python artifacts only.
</ParamField>

<ParamField body="options.virtualenvContent" type="Uint8Array | ArrayBuffer">
  Virtualenv archive contents. Required when `virtualenv` is set.
</ParamField>

<ParamField body="options.mainClass" type="string">
  Main class for the artifact. Jar artifacts only.
</ParamField>

### Return type: `UploadSparkJobResult`

| Field | Type | Description |
| - | - | - |
| `artifact` | `SparkArtifactSummary` | The committed artifact, including its new `version` and `status` |
| `uploadedComponents` | `string[]` | Which parts were uploaded, e.g. `["entrypoint", "pyFiles"]` |

<Note>
  Every upload creates a new artifact version, so re-uploading the same entrypoint never overwrites a previous version in place.
</Note>

***

## `listSparkJobs(options?)`

List your uploaded Spark job artifacts with pagination.

```ts theme={null}
const { artifacts, hasMore } = await client.listSparkJobs();
for (const artifact of artifacts) {
  console.log(artifact.name, artifact.entrypoint, artifact.status);
}

let offset = 0;
const all = [];
while (true) {
  const page = await client.listSparkJobs({ limit: 50, offset });
  all.push(...page.artifacts);
  if (!page.hasMore) break;
  offset += 50;
}
```

### Parameters

<ParamField body="options.limit" type="number" default="20">
  Number of artifacts to return per page.
</ParamField>

<ParamField body="options.offset" type="number" default="0">
  Number of artifacts to skip for pagination.
</ParamField>

### Return type: `ListSparkJobsResult`

| Field | Type | Description |
| - | - | - |
| `artifacts` | `SparkArtifactSummary[]` | Artifacts for the current page |
| `hasMore` | `boolean` | Whether more artifacts are available |

Each `SparkArtifactSummary` has:

| Field | Type | Description |
| - | - | - |
| `id` | `string` | Artifact ID |
| `entrypoint` | `string` | Entrypoint script or class reference |
| `name` | `string` | Artifact name |
| `version` | `number` | Artifact version |
| `language` | `string` | Artifact language |
| `mainClass` | nullable string | Main class for JVM jobs |
| `pyFiles` | nullable string | Python dependencies archive |
| `virtualenv` | nullable string | Virtualenv archive |
| `status` | `string` | Artifact status |
| `createdAt` | nullable string | ISO timestamp when created |
| `updatedAt` | nullable string | ISO timestamp when last updated |

***

## `submitSparkJob(options)`

Submit a Spark job for execution. `cluster` defaults to `oleander`.

For oleander-managed Spark, `entrypoint` is the uploaded script name. For external clusters, `entrypoint` is cluster-specific, such as an S3 URI for EMR Serverless or a Glue job name for Glue.

```ts theme={null}
const { runId } = await client.submitSparkJob({
  namespace: "my-namespace",
  name: "daily-etl",
  entrypoint: "etl_pipeline.py",
  args: ["--date", "2026-03-11"],
  executorNumbers: 4,
});

const run = await client.getRun(runId);
```

### Common options

| Option | Type | Default | Description |
| - | - | - | - |
| `cluster` | `string` | `"oleander"` | Managed cluster or registered cluster name |
| `namespace` | `string` | | Job namespace |
| `name` | `string` | | Job name |
| `entrypoint` | `string` | | Script name, S3 URI, or Glue job name depending on cluster |
| `args` | `string[]` | `[]` | Entrypoint arguments |
| `sparkConf` | `string[]` | `[]` | Spark configuration values |
| `packages` | `string[]` | `[]` | Extra package coordinates |
| `jobTags` | `string[]` | `[]` | Tags applied to the job |
| `runTags` | `string[]` | `[]` | Tags applied to this run |

### Oleander-managed options

| Option | Type | Default | Description |
| - | - | - | - |
| `driverMachineType` | `SparkMachineType` | `spark.1.b` | Driver machine type |
| `executorMachineType` | `SparkMachineType` | `spark.1.b` | Executor machine type |
| `executorNumbers` | `number` | `2` | Number of executors, from 1 to 20 |

### EMR Serverless options

| Option | Type | Description |
| - | - | - |
| `pyFiles` | `string` | Zip archive of Python dependencies |
| `virtualenv` | `string` | Virtualenv archive for Python jobs |
| `mainClass` | `string` | Main class for JVM jobs |
| `executionIamPolicy` | `string` | IAM policy applied to execution |

### Glue options

| Option | Type | Default | Description |
| - | - | - | - |
| `workerType` | `string` | | Glue worker type |
| `numberOfWorkers` | `number` | `1` | Number of Glue workers |
| `enableAutoScaling` | `boolean` | | Enable Glue auto scaling |
| `timeoutMinutes` | `number` | | Timeout in minutes |
| `executionClass` | `string` | | Glue execution class such as `STANDARD` or `FLEX` |
| `executionIamPolicy` | `string` | | IAM policy applied to execution |

<Note>
  For Glue jobs, `args` are converted into key-value pairs. Pass them as alternating entries such as `["--source", "s3://bucket/input", "--target", "s3://bucket/output"]`.
</Note>

### External cluster example

```ts theme={null}
const { runId } = await client.submitSparkJob({
  cluster: "emr-prod",
  namespace: "finance",
  name: "daily-etl",
  entrypoint: "s3://my-bucket/jobs/etl_pipeline.py",
  args: ["--date", "2026-03-11"],
  pyFiles: "s3://my-bucket/jobs/deps.zip",
  packages: ["org.example:my-lib:1.0.0"],
});
```

### Machine types

The `SparkMachineType` enum covers compute-optimized (`c`), balanced (`b`), and memory-optimized (`m`) options:

| Type | vCPUs | Category |
| - | - | - |
| `spark.1.c` / `spark.1.b` / `spark.1.m` | 1 | Compute / Balanced / Memory |
| `spark.2.c` / `spark.2.b` / `spark.2.m` | 2 | Compute / Balanced / Memory |
| `spark.4.c` / `spark.4.b` / `spark.4.m` | 4 | Compute / Balanced / Memory |
| `spark.8.c` / `spark.8.b` / `spark.8.m` | 8 | Compute / Balanced / Memory |
| `spark.16.c` / `spark.16.b` / `spark.16.m` | 16 | Compute / Balanced / Memory |

***

## `submitSparkJobAndWait(options)`

Submit a Spark job and poll until it reaches a terminal state (`COMPLETE`, `FAIL`, or `ABORT`). Throws an error if the timeout is exceeded.

```ts theme={null}
const { runId, state, run } = await client.submitSparkJobAndWait({
  namespace: "my-namespace",
  name: "daily-etl",
  entrypoint: "etl_pipeline.py",
  pollIntervalMs: 5000,
  timeoutMs: 300000,
});

if (state === "COMPLETE") {
  const elapsed = run.duration;
  // proceed with downstream work ...
} else {
  throw new Error(`Run ${runId} ended with state: ${state}`);
}
```

Accepts all `submitSparkJob` options plus:

<ParamField body="options.pollIntervalMs" type="number" default="10000">
  Milliseconds between status polls.
</ParamField>

<ParamField body="options.timeoutMs" type="number" default="600000">
  Maximum time to wait in milliseconds before throwing a timeout error.
</ParamField>

***

## `getRun(runId)`

Get the current status of a run. Use this to poll a job submitted with `submitSparkJob()`.

```ts theme={null}
const run = await client.getRun(runId);

if (run.state === "COMPLETE") {
  const duration = run.duration;
  const jobName = run.job.name;
  // handle completion ...
} else if (run.state === "FAIL") {
  const error = run.error;
  // handle failure ...
}
```

### Return type: `RunResponse`

| Field | Type | Description |
| - | - | - |
| `id` | `string` | Run ID |
| `state` | nullable string | Current state |
| `started_at` | nullable string | ISO timestamp when the run started |
| `queued_at` | nullable string | ISO timestamp when the run was queued |
| `scheduled_at` | nullable string | ISO timestamp when the run was scheduled |
| `ended_at` | nullable string | ISO timestamp when the run ended |
| `duration` | nullable number | Run duration in seconds |
| `error` | unknown | Error details if the run failed |
| `tags` | array | Array of `{ key, value, source }` objects |
| `job` | object | Job info with `id`, `name`, `namespace` |
| `pipeline` | object | Pipeline info with `id`, `name`, `namespace` |


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.