TimeSeries Code for Critical Systems
a Time-series R&D Agent for Critical Infrastructures, Facilities, and Equipments
Our research on Critical AIOps focuses on the shared time-series analysis problems that arise in the operation and maintenance of critical systems. We are developing TimeSeries Code, an R&D agent for time-series intelligence, to help users turn local needs into tested Time Series Agents that can run in their own operating environment. The development process continues as operational evidence accumulates and users clarify what they need.
Project materials
Presentation slides (PDF)
The architecture and an end-to-end development example. 25 bilingual slides.
Both videos use the bilingual slides, with captions in the spoken language and authorized AI-generated narration.
Research motivation
We use critical systems to refer to critical infrastructure and major industrial and scientific equipment and facilities. These systems produce large volumes of continuous, often real-time numerical observations, from physical sensors and operational telemetry. Temperature, vibration, power and performance measurements record how a system behaves over time. Our research on time-series intelligence asks how to use these observations to perceive system state, reason about what is happening, predict what may happen next, and support operational decisions. Data centers, power systems, aircraft, manufacturing equipment and scientific facilities all offer important settings for this work.
Abundant data, however, do not by themselves make these capabilities reliable. A fundamental difficulty is the gap between model development and on-premise deployment. Algorithms and models that work well on training and evaluation datasets can perform poorly at a user’s site, where data distributions differ. Equipment, workloads, control settings and operating practices vary across sites. Even two installations of the same type of equipment need not produce the same patterns.
Data quality adds another difficulty. Local observations are often less complete and less consistent than the curated data used during development and evaluation. Sensors may be noisy or drift; measurements may be missing, sampled at different rates or recorded with misaligned timestamps. Large archives of routine operation may contain very few labeled failures. Interpreting them also requires context that is not captured in the numbers alone: a rise in a generating unit’s temperature, for example, could reflect a change in load rather than a developing fault.
The problem continues after deployment. Maintenance, equipment aging, software updates and changing operating conditions can alter what counts as normal. A detector that was useful when commissioned may later generate excessive alarms or miss important changes. Good performance on a fixed dataset is therefore not enough; the method must remain useful as the system and its use evolve.
Meeting this need involves more than choosing a model. Data preparation, domain knowledge, algorithm development, testing and deployment must work together under local constraints on latency, computing resources, privacy and safety. In many of the environments we target, data and execution records cannot leave the site. Developing and validating a useful solution must therefore be possible locally, without routinely sending data elsewhere for analysis or retraining.
This is why we are developing TimeSeries Code as an R&D agent. We want to make it easier to turn a local operational need into a tested, usable method, and to revise that method as evidence accumulates. Some requirements only become clear after deployment. Supporting that repeated development process is as important as producing the first algorithm.
From a development request to Time Series Agents
The design distinguishes development from operation. TimeSeries Code is the R&D agent. It helps develop, validate and improve Time Series Agents for particular systems and tasks, such as anomaly detection, forecasting, diagnosis, or health and performance analysis. Both run on an Agent Platform such as OpenTrek. Sharing a platform does not require the same process, model instance or permissions.
ChatTS is the shared interaction workspace. Users bring requirements, available domain knowledge and feedback here, explore observations, inspect results and annotate evidence. Topology, sensor meanings, operating constraints and acceptance criteria can be supplied and clarified over time. Auxiliary views can be generated when a task calls for them, rather than requiring a fixed interface for every investigation. Interaction records and annotations remain available for later development.
The architecture distinguishes four interactions. User requirements and user feedback enter ChatTS separately. Operational data enter the delivered agents, which produce agent outputs such as findings, forecasts and recommendations. Results can also be presented and discussed through ChatTS. Development uses authorized local observations and evidence as needed; these paths distinguish responsibilities, not separate pools of information.
A shared foundation for development and operation
ChatTS FM is the shared time-series multimodal foundation model, supporting language, vision, audio and structured knowledge alongside time series. It is distinct from the ChatTS interface. We describe three capabilities within this foundation. EvidenceTS asks what an observation supports or contradicts, with attention to provenance, uncertainty, event time, observation time, availability and admissibility. ReasonTS draws conclusions from that evidence and can request more observations or experiments. ComposeTS writes, tests and debugs code and workflows, using algorithms in LibraryTS and reusable skills in SkillTS. The platform executes the tools and code; the model does not execute them inside its weights.
Both development and operational agents use this foundation and these assets. An operational agent need not be confined to a fixed pipeline: it can analyze an unfamiliar situation, reason from additional evidence and produce code on demand, within its authority. Such a temporary analysis does not automatically become a persistent update to the deployed agent.
Three modules within TimeSeries Code
AutoResearchTS organizes task-focused development: propose an approach, implement it, run experiments, evaluate the results, and revise it. Algorithm selection and parameter search are part of this work, but not its limit. It calls ComposeTS for implementation and GymTS for experimental evidence, works within explicit budgets and stopping criteria, and organizes review and approved delivery. A straightforward correction can follow a short implementation-and-test path; it need not trigger a lengthy research cycle.
GymTS provides deployable experimentation and evaluation capabilities at the user’s site. EnvironmentTS integrates local history, traces and annotations with domain simulators, testbeds and replay environments. It provides interaction and reward interfaces and preserves results and execution context. AutoBenchTS supports automated reproduction, acceptance, regression and safety evaluation, including benchmark construction and maintenance. Local tasks, operating conditions and acceptance criteria determine the tests. This does not require us to build every industry’s simulator or testbed ourselves.
EvolveTS organizes recursive self-improvement of the delivered agents. It combines operational outcomes, traces and relevant feedback with current skills, policies and versions, experiments from GymTS, and review conclusions and version status from AutoResearchTS. It identifies an improvement task, coordinates development and follows the update back into operation to check whether it helped.
Skill-level improvement revises instructions, code and workflows through AutoResearchTS and ComposeTS, followed by testing and review. It does not require local model training. Agentic RL is optional: suitable models, authorized training data and reward signals, training resources and approval must all be available. A trace records what happened; it is not itself a reward. A resulting model or adapter is a candidate for the target agent, not an automatic replacement for the shared ChatTS FM.
Human oversight and authorization apply across all three modules. Users can inspect evidence and intervene. Benchmark inclusion, asset releases and persistent agent updates require the appropriate review; candidates cannot approve their own deployment. Releasing a version in LibraryTS or SkillTS is also separate from enabling it for a particular agent. This keeps local knowledge and working versions from being silently overwritten.
The loop shown here concerns work at the user’s site. Local data and traces stay there by default. If model training is not possible locally, the design does not assume that traces can be sent elsewhere for training. Skill and workflow improvements remain useful without that step. Developing the project’s general models and R&D capabilities is separate from this operational improvement loop.
Open the high-resolution architecture diagram
Application domains
We organize this research around seven domains, using concrete systems to guide the development questions. The examples below describe the research scope, not completed deployments.
- Information infrastructure: data centers, computing networks, communications, and navigation and timing facilities.
- Energy systems: power generation and storage, transmission and distribution, electricity use and control, oil and gas, heat supply, coal facilities, and hydrogen production, storage and transport.
- Transportation, aviation and aerospace: road and rail vehicles, ships and submarines, aircraft and engines, drones, spacecraft, satellite platforms, and supporting facilities and networks.
- Major industrial equipment and production facilities: lithography systems, advanced machine tools, equipment test benches, tunnel boring machines, and industrial production facilities.
- Major scientific facilities: research wind tunnels, propulsion testing facilities, fusion facilities, accelerators, light and neutron sources, and astronomical observatories.
- Environmental sensing and monitoring: weather radar, ocean sensing and sonar, and networks for weather, ocean, earthquake, air quality, water quality and geological safety monitoring.
- Water infrastructure and urban utility networks: reservoirs, dams, pumping stations, water supply and treatment, drainage, and underground utility networks.
Open the high-resolution application domains diagram
An example: developing an online anomaly detector
Consider an operations team that asks: “Develop an online anomaly detection system using our KPI data.” Through ChatTS, the team clarifies how results will be used, supplies available system knowledge and agrees on acceptance criteria. AutoResearchTS uses ChatTS FM and ComposeTS to develop the workflow, draws on LibraryTS and SkillTS, and works with GymTS to compare and test alternatives. The chosen version is frozen before independent acceptance on held-out data. After review and authorization, it is delivered as a time series agent on the platform.
Suppose the team later receives too many reports at night. A request to “stop reporting at night” still needs discussion. The team may want to retain detection records while pausing routine notifications. Once ChatTS clarifies the requirement, EvolveTS combines it with the current version and operational evidence to define an improvement task. AutoResearchTS calls ComposeTS to revise the implementation and compares alternatives where needed. GymTS provides replay, acceptance and regression tests. Only after review and approval is the target agent updated. Reporting policy and detection quality need separate evaluation.
This is an illustrative example, but it explains why operational feedback belongs in the development process. Some requirements only become clear when people use the system. We want TimeSeries Code to support that continuing work: turning observations into useful evidence, developing methods that can run locally, and testing each improvement before it is relied on.