14 Operator Improvements and Best Practices
This chapter covers advanced techniques for improving operator quality, reliability, and maintainability. You’ll learn essential practices for logging, error handling, testing, and optimization that ensure your operators work robustly in production environments.
By the end of this chapter, you will be able to: - Implement comprehensive logging and debugging strategies - Build robust error handling and input validation - Create comprehensive test suites for operators - Apply performance optimization techniques - Follow best practices for production-ready operators
Before proceeding, ensure you’ve completed: - Basic Implementation chapter for core operator concepts - Advanced Features chapter for complex functionality - Data Input and Output Patterns for data handling
14.1 Logging and Debugging
Effective logging is essential for monitoring operator behavior and diagnosing issues in production environments.
14.1.1 Basic Logging
Implement logging for production operators:
ctx$log("Your message.")ctx.log("Your message.")- Log key milestones: Start/end of major operations
- Include data metrics: Row counts, processing times, memory usage
- Log parameter values: Help reproduce issues with specific inputs
- Use structured formats: Enable easier log parsing and analysis
- Avoid logging sensitive data: Protect user privacy and security
14.2 Error Handling
Implement comprehensive error handling that provides helpful feedback to users:
# Comprehensive error handling with user-friendly messages
robust_operator <- function(ctx) {
tryCatch({
# Validate inputs first
validate_inputs(ctx)
# Main processing with progress logging
ctx$log(paste("[INFO]", Sys.time(), "- Beginning data analysis"))
# Check for edge cases
data <- ctx$select(.ri, .ci, .y)
if (any(is.infinite(data$.y))) {
ctx$log(paste("[WARNING]", Sys.time(), "- Infinite values detected, removing them"))
data <- data[is.finite(data$.y), ]
}
if (nrow(data) == 0) {
stop("No valid data remaining after cleaning")
}
# Perform analysis
result <- perform_analysis(data)
ctx$log(paste("[INFO]", Sys.time(), "- Analysis completed successfully"))
return(result)
}, error = function(e) {
# Log the technical error
ctx$log(paste("[ERROR]", Sys.time(), "- Technical error:", e$message))
# Provide user-friendly error message
if (grepl("projection.*required", e$message)) {
stop("Please ensure you have dragged the required data columns to the appropriate axes.")
} else if (grepl("data points required", e$message)) {
stop("This analysis requires at least 3 data points. Please check your data selection.")
} else if (grepl("values must vary", e$message)) {
stop("The data values do not vary enough for this analysis. Please check your input data.")
} else {
stop(paste("An error occurred during analysis:", e$message))
}
})
}def robust_operator(tercen_ctx):
"""Operator with comprehensive error handling"""
try:
# Validate inputs first
validate_inputs(tercen_ctx)
# Main processing with progress logging
tercen_ctx.log(f"[INFO] {datetime.now()} - Beginning data analysis")
# Check for edge cases
df = tercen_ctx.select(['.ri', '.ci', '.y'], df_lib="polars")
# Handle infinite values
infinite_count = df.filter(pl.col('.y').is_infinite()).height
if infinite_count > 0:
tercen_ctx.log(f"[WARNING] {datetime.now()} - Infinite values detected ({infinite_count}), removing them")
df = df.filter(pl.col('.y').is_finite())
if len(df) == 0:
raise ValueError("No valid data remaining after cleaning")
# Perform analysis
result = perform_analysis(df)
tercen_ctx.log(f"[INFO] {datetime.now()} - Analysis completed successfully")
return result
except ValueError as ve:
tercen_ctx.log(f"[ERROR] {datetime.now()} - Validation error: {str(ve)}")
# Provide user-friendly error messages
if "projection" in str(ve) and "required" in str(ve):
raise ValueError("Please ensure you have dragged the required data columns to the appropriate axes.")
elif "data points required" in str(ve):
raise ValueError("This analysis requires at least 3 data points. Please check your data selection.")
elif "values must vary" in str(ve):
raise ValueError("The data values do not vary enough for this analysis. Please check your input data.")
else:
raise ValueError(f"An error occurred during analysis: {str(ve)}")
except Exception as e:
tercen_ctx.log(f"[ERROR] {datetime.now()} - Unexpected error: {str(e)}")
raise ValueError(f"An unexpected error occurred. Please check your data and try again. Error: {str(e)}")- Validate early: Check inputs before expensive computations
- Fail gracefully: Provide clear, actionable error messages
- Log technical details: Help with debugging while keeping user messages simple
- Handle edge cases: Account for missing data, infinite values, empty datasets
- Test error scenarios: Ensure error handling works as expected
14.3 Memory: tell Tercen what your operator needs
When Tercen runs your operator it puts the container under a hard memory limit. That limit is decided before the run starts, from the first of these that applies:
- a user-set
ramon the step — always respected verbatim; - the operator’s own
memory_model.json, evaluated against this run’s data; - a generic platform estimate when the operator ships no model.
Only the second option knows anything about your operator. An operator whose real usage sits near or above whatever it is given will fail — and it does so with two different signatures, which is what makes this confusing to diagnose:
Failed to run operator, please increase memory resources. (exit code 137)
is the operator process itself being killed at the limit. But when the operator shells out — an R or Python wrapper invoking a compiled binary (system2(), subprocess) — the kernel kills the child, the wrapper only sees “the thing I called failed”, and the platform reports:
Failed to run operator, exit code = 1
with stderr that simply stops mid-progress — a progress bar frozen at 11%, a half-written log, no error message. There is no “memory” in that text anywhere. Treat truncated output with no real error as an OOM until proven otherwise; the fix is identical.
Near-ceiling runs also fail intermittently: the run sits exactly on its limit, and whether the kernel can reclaim enough cache decides its fate. Same data, different outcome.
14.3.1 The quick fix (no release needed)
On the failing workflow step, add the environment setting ram with the size you want — 2 GB, 512 MB, 16 GB — then reset the step and rerun. A user-defined value is always respected verbatim, so pick generously: over-booking only means the task waits for a bigger worker, while under-booking fails the run.
Two things to know about the field. A value with no unit is read as bytes, so 2 books two bytes rather than 2 GB. And human units need tercen 1.0.23 or newer; on older instances write the byte count (2000000000).
14.3.2 The durable fix: memory_model.json
Ship a memory_model.json at the repository root. Tercen reads it at install/update time and books memory per run based on the projected data:
{
"kind": "OperatorEstimateModel",
"intercept": 1.0,
"offset": 500.0,
"features": [
{ "kind": "Feature", "name": "n_main", "coefficient": 0.02, "exponent": 1.0 }
]
}The estimate, in MB, is intercept × Π(coefficient × feature^exponent) + 1.5 × offset — the example books 750 MB plus 0.02 MB per projected cell.
Available features:
| Feature | Value |
|---|---|
n_main |
number of projected crosstab cells |
n_cols |
rows of the column table |
n_rows |
rows of the row table |
settings.<Name> |
numeric value of an operator setting |
a_||_b |
interaction: product of two features |
Prefer the schema features. A settings.* feature whose setting is left blank evaluates to 0, which can silently disable the whole model (an infinite or zero estimate is discarded).
14.3.3 How the formula is evaluated — worked example
The model is multiplicative in the features, plus a constant:
estimate_MB = intercept × ∏ ( coefficient_i × feature_i ^ exponent_i ) + 1.5 × offset
Take the model above and a crosstab of 200 kinases × 12 samples, so n_main = 2 400 cells:
growth = 1.0 × (0.02 × 2400^1.0) = 48 MB
constant = 1.5 × 500 = 750 MB
booked = 48 + 750 = 798 MB
At 100 000 cells the same model books 750 + 2 000 = 2 750 MB — the booking follows the data instead of being a fixed number.
Two structural consequences of this shape are worth internalising:
offsetis the additive floor (times 1.5). It represents everything that exists before your data does: the R or Python runtime, loaded libraries, reference databases read into memory. For many operators this constant is the whole story.interceptand the features form the growth term. With several features the terms multiply, so two linear features model ann × msurface, not a sum. If you want additive behaviour, put it inoffsetand keep one feature.
14.3.4 Three real models, three shapes
Constant-dominated — pamgene/cs_uka_operator. Measured production runs peaked at 525–592 MB whether the crosstab had 600 cells or 12 000: the UKA reference database and the permutation machinery dominate; the projected data is a rounding error. The model is therefore mostly floor, with a token slope as insurance:
{
"kind": "OperatorEstimateModel",
"intercept": 1.0,
"offset": 500.0,
"features": [
{ "kind": "Feature", "name": "n_main", "coefficient": 0.02, "exponent": 1.0 }
]
}750 MB + 0.02 MB/cell — fitted against 203 recorded runs, covering every one with 1.3–1.9× headroom. (Before this model existed, the operator got the 500 MiB default: every run sat exactly on its ceiling, and the intermittent exit-137 reports rolled in for months.)
Linear in the data — tercen/plot_operator. A plot’s memory scales with the points drawn, on top of an R baseline:
{
"kind": "OperatorEstimateModel",
"intercept": 0.0008,
"offset": 1000.0,
"features": [
{ "kind": "Feature", "name": "n_main", "coefficient": 1.0, "exponent": 1.0 }
]
}1 500 MB + 0.0008 MB/point: a 1M-point scatter books ~2 300 MB, a 10M-point one ~9 500 MB. Note the coefficient/intercept split is arbitrary — only their product matters (0.0008 × 1.0 ≡ 1.0 × 0.0008).
Fitted power-law — tercen/fast_tSNE_operator. t-SNE’s footprint grows sub-linearly with observations and weakly with dimensions, so the fit found fractional exponents over two interacting features:
{
"kind": "OperatorEstimateModel",
"intercept": 0.0076,
"offset": 238.5,
"features": [
{ "kind": "Feature", "name": "n_cols", "coefficient": 1.0, "exponent": 0.9189 },
{ "kind": "Feature", "name": "n_rows", "coefficient": 1.0, "exponent": 0.3136 }
]
}For 100 000 cells × 20 markers: 0.0076 × 100000^0.9189 × 20^0.3136 + 1.5 × 238.5 ≈ 0.0076 × 39 000 × 2.56 + 358 ≈ 1 117 MB. The two features multiply — that is the interaction between how many observations you embed and how many dimensions each carries.
14.3.5 Fit it from real runs, don’t guess
Tercen records ground truth for you: every operator run stores its actual memory peak on the task document — stats_d_actual_ram_peak (total cgroup peak, includes reclaimable file cache) and stats_d_actual_ram_peak_anon (bytes the operator actually allocated; the better fitting target). The task also stores stats_d_qt_nrows, stats_d_col_nrows, stats_d_row_nrows — the feature values. The recipe:
- Collect the peaks and dimensions from historical runs (the more varied the sizes, the better the fit).
- Look at the scatter first. Flat cloud → constant-dominated (fit the floor, token slope). Straight line → linear. Bends on a log-log plot → power law with fractional exponent.
- Fit, then verify the safety property: the model must predict at or above every observed peak, with 1.3–2× headroom on the median. A memory model that under-books reintroduces the exact failure it exists to prevent; over-booking merely schedules conservatively.
- Check the extrapolation at 10× your largest observed run — the number should be defensible, because one day someone will run it.
tercen/plot_operator ships an estimate-memory.yml GitHub workflow that automates this fit — a good starting point if you want it continuous.
completedDate, never lastModifiedDate
Read a task’s duration as
completedDate - createdDate
lastModifiedDate is not when the work finished — it is when the document was last written, and plenty writes to it afterwards: the GC sweep marking the task, a web app’s teardown when the user closes it, a clone of the task, a state correction. Subtracting it yields a number that is silently, wildly wrong.
Two real examples from the same instance, both on tasks that in fact took about nine seconds:
| Task | lastModified − created |
completedDate − created |
What touched it |
|---|---|---|---|
| An image operator | 523.8 s | 8.8 s | GC sweep, eight minutes later |
| A web app step | 77 983 s | — | the app’s teardown, a day later |
The failure mode is the dangerous kind: the number looks plausible enough to report. A 523-second reading was very nearly written up as a 25× performance regression before the scheduler log showed the task had gone running → done nine seconds after it started.
Use completedDate for any timing comparison — A/B tests, before-and-after measurements, and the runtime half of a resource model.
Failed (exit-137) runs are censored observations: they tell you the true need exceeded that run’s limit, not what it was. Fit on successful runs, then check the model clears every failed run’s limit at its recorded dimensions.
14.3.6 Verify the model against a real run
Once a model ships, one successful run tells you whether it is honest. Read the two numbers back off the task and divide:
booked = stats_d_estimated_ram (the hard limit the model asked for)
actual = stats_d_actual_ram_peak_anon (what the run really allocated)
headroom = booked / actual
A production example: tercen/fast_tSNE_operator 0.2.1 on 695 052 observations × 8 variables booked 3.95 GB and peaked at 2.61 GB — 1.51× headroom, comfortably inside the 1.3–2× target, on a dataset ~7× larger than anything in the original fitting sweep. That is what a healthy model looks like at unfamiliar scale.
The same dataset run through an earlier release of that operator, which shipped no model at all, was booked 2.13 GB by the generic estimate and died part-way through with exit code = 1 and a frozen progress bar. Same code, same data — the model was the entire difference.
14.3.7 CPU is booked too, and it gates your threads
Memory gets the attention, but the step also books CPUs (cpu), and the booking confines your threads. Code written to use every available core — nthreads = 0 in a FIt-SNE wrapper, mclapply, a Rust operator on Rayon — still runs single-threaded when the task was booked one CPU. Nothing errors; it is simply several times slower.
Measured on the same t-SNE dataset: 6 minutes with cpu = 4 versus 20 minutes at the default cpu = 1. If your operator scales with cores, say so — declare a CPU need or document the setting — otherwise users get the slow path by default and conclude the operator is slow.
14.3.8 Memory your model cannot see: file inputs
memory_model.json features describe the crosstab: projected cells, column-table rows, row-table rows. That is the right vocabulary for an operator whose memory follows its table — and the wrong one for an operator whose memory follows something else entirely.
The clearest case is a file input. An operator that takes a documentId and downloads an archive — instrument exports, image bundles, reference data — has memory driven by bytes on the wire, while the model can only see “one row, one column”. Ten rows of image metadata and a 2 GB archive look identical to every available feature. No fit, however careful, can close that gap.
The fix is not a better model, it is not holding the file in memory:
Download to a temp file, read from the file, and process entries one at a time. Peak memory then depends on your buffer size, not on the input size — and the flat offset in your model becomes truthful for every dataset, however large.
tercen/ps12image_rust_operator is a worked before/after. It originally accumulated the download into a Vec<u8>, kept that buffer resident while unpacking, and allocated a second buffer per archive entry — peak memory ≈ archive size + largest entry. Fine for a 100 MB test export; a 1–2 GB instrument export would blow straight past a 400 MB booking that, by construction, could never anticipate it.
The 0.1.2 rewrite streams instead: gRPC chunks are written straight to a temp file, the ZIP reader works from that file (it needs only Read + Seek), and each entry is extracted with a fixed-size copy rather than a per-entry buffer:
// before — peak grows with the archive
let bytes = stream_file_bytes(ctx, doc_id).await?; // whole archive in RAM
let mut archive = ZipArchive::new(Cursor::new(&bytes))?;
let mut data = Vec::with_capacity(entry.size() as usize);
entry.read_to_end(&mut data)?; // + largest entry
out.write_all(&data)?;
// after — peak is constant
stream_file_to(ctx, doc_id, &archive_path).await?; // chunks -> disk
let mut archive = ZipArchive::new(File::open(&archive_path)?)?;
std::io::copy(&mut entry, &mut out)?; // fixed bufferPeak dropped from O(archive) to O(64 KB), with byte-identical output — the golden unit test proved the change was purely mechanical. About thirty lines, and the operator became indifferent to input size forever.
The general principle: before fitting a model, ask what actually drives the memory. If the answer is not visible in the crosstab dimensions, fix the code rather than the model.
Because the file is read from git at install/update time, shipping or changing it is a normal release (bump the container pin, tag) — but it touches no code: results, the unit test, and the image content are unchanged.
14.4 Testing and Validation
Comprehensive testing ensures your operator works correctly across different data scenarios and edge cases.
Tercen supports two main testing frameworks: 1. Unit Tests: Simple data files with expected input/output and test specifications 2. Integration Tests: Actual Tercen workflows triggered to perform computations
14.4.1 Unit Test Structure
Create a tests directory in your operator repository with the following structure:
tests/
├── input.csv # Sample input data
├── output.csv # Expected output data
├── test.json # Test configuration
For multiple test scenarios, use numbered files: - test_1.json, test_2.json for different parameter settings - input_1.csv, input_2.csv for different data scenarios
14.4.2 Creating Comprehensive Test Data
Design test cases that cover various scenarios:
| Test Scenario | Purpose | Example Data |
|---|---|---|
| Normal Case | Standard operation | Regular numeric data with good distribution |
| Edge Cases | Boundary conditions | Minimum data points, extreme values |
| Error Cases | Invalid inputs | Missing data, wrong data types |
14.4.3 Test Configuration File
Create a tests/test.json in OperatorUnitTest format. This is what the release install-check and the library gate execute; R/testthat scripts do not count for the gate. A complete, battle-tested reference: pamgene/cs_uka_operator/tests/.
{
"kind": "OperatorUnitTest",
"name": "default_params",
"namespace": "ds0",
"inputDataUri": "test_in.csv",
"outputDataUri": ["test_out_1.csv", "test_out_2.csv"],
"columns": ["Supergroup"],
"rows": ["ID"],
"colors": [],
"labels": ["colSeq"],
"yAxis": "value",
"xAxis": "grp",
"equalityMethod": "R2",
"r2": 0.99,
"propertyValues": [
{"kind": "PropertyValue", "name": "UkaDbVersion", "value": "0.6"},
{"kind": "PropertyValue", "name": "Seed", "value": "42.0"}
]
}Rules that each come from a real gate failure:
- Pin every operator property via
propertyValues(a list ofPropertyValueobjects — values as strings). A fully pinned test never breaks when a bundled database or a default changes; it only changes when you change it. (Note: the field ispropertyValues, notproperties.) equalityMethod: "R2"for computed floats. The defaultequalscompares doubles near-exactly and fails on floating-point noise (task.test.operator.bad.value). Reserveequalsfor exact integer/string outputs.- One entry in
outputDataUriper output relation — a result joined to the crosstab typically yields 2+ relations; a count mismatch fails withtask.test.operator.bad.nRelations. - Pin column types with
.schemasidecars: for each expected CSV also commit<name>.csv.schema— the result relation’sTableSchemaJSON (columns with name+type,id/revstripped). Without it the gate re-infers types from CSV text and fails withtask.test.operator.bad.column.type — expected double found int32on columns of whole numbers. Write doubles type-faithfully too (58.0, not58).
14.4.4 Determinism first: the Seed rule
A golden test is impossible if two runs differ. Any operator using randomness (permutations, sampling, clustering) must seed the RNG, best as a visible setting:
seed <- ctx$op.value("Seed", as.double, 42)
if (seed >= 0) set.seed(seed)with a matching Seed DoubleProperty (default 42; -1 = unseeded). Acceptance test: run twice, outputs byte-identical.
14.4.5 Generating and regenerating the expected output
The expected CSVs are golden files from a real run of the exact code being released — they cannot be hand-written. Generate them on Tercen Studio (see the appendix), never on a production instance: install the operator from the release branch, import the small input, run the step with the projection and pinned settings from test.json, run it twice (determinism), export each output relation and its schema.
When a bundled database or intended behaviour changes, the release install-check goes red — that is the alarm: re-run the same input with the new settings, replace the expected files, update the changed propertyValues, bump the version, release. Ship this recipe in the operator repo itself as a CLAUDE.md (the create-operator skill provides a template), so any future maintenance session knows the drill.
This completes our comprehensive guide to Tercen operator development. You now have all the tools and knowledge needed to create robust, efficient, and user-friendly operators that extend Tercen’s analytical capabilities!