diff --git a/feature_integration_tests/README.md b/feature_integration_tests/README.md index 6f9ced5daf9..33f20bdf4fb 100644 --- a/feature_integration_tests/README.md +++ b/feature_integration_tests/README.md @@ -48,6 +48,97 @@ bazel run //feature_integration_tests/test_scenarios/rust:rust_test_scenarios -- bazel test --config=linux-x86_64 //feature_integration_tests/test_cases:fit --test_output=streamed ``` +To run the lifecycle scenario-stub tests directly with `pytest` and build the scenario binaries on demand: + +```sh +python3 -m pytest feature_integration_tests/test_cases/tests/lifecycle/test_conditional_launching_scenario.py \ + --build-scenarios -m rust -q -v + +python3 -m pytest feature_integration_tests/test_cases/tests/lifecycle/test_conditional_launching_scenario.py \ + --build-scenarios -m cpp -q -v +``` + +The daemon-driven lifecycle tests (`test_conditional_launching.py`, `test_process_launching_with_daemon.py`, +`test_retry_exhaustion.py`) resolve `launch_manager`, the supervised apps and the config tools from the +`FIT_*_PATH` variables set by their Bazel targets, and carry no `rust`/`cpp` marker, so run them via +`bazel test //feature_integration_tests/test_cases:fit_lifecycle_daemon` / `:fit_lifecycle_retries` +rather than plain `pytest -m rust|cpp` (which would deselect them). + +#### Sandbox uid/gid and scheduling-policy tests + +Some lifecycle daemon tests (`test_launched_process_uid_gid_matches_config_when_applied`, +`test_launched_process_scheduling_matches_config_when_applied`, +`test_scheduling_policy_is_non_default_and_applied`) verify that `launch_manager` +applies the sandbox `uid`/`gid` and scheduling policy from +`feature_integration_tests/configs/lifecycle_daemon_config.json`. This requires granting +`launch_manager` the `cap_setuid,cap_setgid,cap_sys_nice,cap_kill` file capabilities via `setcap`, which +in turn requires `CAP_SETFCAP` — not available to a non-root test runner by default, so these +tests opt in via the `FIT_ENABLE_SETCAP` env var (backed by a passwordless sudoers rule scoped +to the `setcap` binary, e.g. ` ALL=(root) NOPASSWD: /usr/sbin/setcap`, with no trailing +arguments pinned — the target path is a fresh `tmp_path` on every run) and skip otherwise. + +Under `bazel test`, undeclared env vars like `FIT_ENABLE_SETCAP` only reach the test process when +passed via `--test_env` (not `--action_env`, which only affects build actions). Two variants of +the full suite are relevant: + +```sh +# Default: matches CI/CD exactly (sandboxed, no FIT_ENABLE_SETCAP) — the three capability tests skip. +bazel test --config=linux-x86_64 --nocache_test_results //feature_integration_tests/test_cases:fit \ + --test_output=all --test_arg=-rs --test_verbose_timeout_warnings + +# Local verification: also exercises the uid/gid and scheduling-policy grants instead of skipping. +bazel test --config=linux-x86_64 --nocache_test_results //feature_integration_tests/test_cases:fit \ + --spawn_strategy=local --test_env=FIT_ENABLE_SETCAP=1 \ + --test_output=all --test_arg=-rs --test_verbose_timeout_warnings +``` + +Flag rationale (shared by both commands unless noted): + +- `--nocache_test_results`: forces re-execution instead of replaying a cached PASS/SKIP, so a + fresh `setcap` attempt is made every time. +- `--test_output=all`: prints full stdout/stderr for every test, not just failures, so the + `sandbox_privileged_reason` and pytest skip-reason diagnostics are visible. +- `--test_arg=-rs`: forwards pytest's `-rs` flag, which prints the reason for every `SKIPPED` + test instead of just `SKIPPED` with no context. +- `--test_verbose_timeout_warnings`: warns when a test's actual runtime is far from its declared + `timeout`/`size`, useful for right-sizing `fit_lifecycle_daemon`'s `timeout = "long"`. +- `--spawn_strategy=local` (local-verification command only): runs the test action directly on + the host instead of inside Bazel's `linux-sandbox`. The sandbox sets `PR_SET_NO_NEW_PRIVS`, + which makes `setuid` (`sudo`) and file capabilities (`setcap`) inert at exec time even with a + correctly configured host — the grant is applied but silently dropped when the supervised + binary later executes. Only unsandboxed execution lets the grant persist. +- `--test_env=FIT_ENABLE_SETCAP=1` (local-verification command only): opts the test process into + the `sudo -n setcap` attempt; without it these tests always take the plain, non-sudo `setcap` + path and skip on a non-root runner. `bazel run` inherits the shell environment directly, so + `export FIT_ENABLE_SETCAP=1` beforehand is sufficient there instead of `--test_env`. + +#### Tests skipped in CI/CD + +The GitHub Actions runners (`ubuntu-latest`, see `.github/workflows/build_and_test_linux.yml`) run +`bazel test` sandboxed (default `linux-sandbox` strategy) and do not set `FIT_ENABLE_SETCAP` or +provision a passwordless `sudo setcap` rule. As a result, the following subtests in +`fit_lifecycle_daemon` always skip in CI, for both the `rust` and `cpp` supervised-app variants: + +- `test_process_launching_with_daemon.py::TestProcessLaunchingWithDaemon::test_launched_process_uid_gid_matches_config_when_applied[rust|cpp]` +- `test_process_launching_with_daemon.py::TestProcessLaunchingWithDaemon::test_launched_process_scheduling_matches_config_when_applied[rust|cpp]` +- `test_process_launching_with_daemon.py::TestProcessLaunchingWithDaemon::test_scheduling_policy_is_non_default_and_applied[rust|cpp]` + +Reason: all three depend on `launch_manager` successfully gaining `cap_setuid,cap_setgid,cap_sys_nice` +via `setcap` (see `daemon_helpers._grant_sandbox_capabilities`), which fails in CI for two +independent reasons, either sufficient on its own: + +1. **Sandboxed execution**: `linux-sandbox` sets `PR_SET_NO_NEW_PRIVS`, making any `setuid`/file-capability + escalation inert at exec time, so even a successful `setcap` call has no effect on the process + that actually runs. +2. **No opt-in / no sudoers rule**: `FIT_ENABLE_SETCAP` is not set in the CI workflow, so the tests + never attempt the `sudo -n setcap` path; and the CI runner has no passwordless sudoers entry for + `setcap` regardless. + +This is by design: `_grant_sandbox_capabilities` degrades gracefully (never raises) and the three +capability-dependent subtests self-skip with a diagnostic reason instead of failing the build. All +other subtests in `fit_lifecycle_daemon` only check same-uid process behavior and require no +privilege escalation, so they run and pass normally in CI. + ### ITF Tests (QEMU-based) ITF tests run on a QEMU target and require the `itf-qnx-x86_64` config: @@ -62,6 +153,7 @@ Test scenarios can be listed and run directly for debugging: ```sh bazel run //feature_integration_tests/test_scenarios/rust:rust_test_scenarios -- --list-scenarios +bazel run //feature_integration_tests/test_scenarios/rust:rust_lifecycle_test_scenarios -- --list-scenarios bazel run --config=linux-x86_64 //feature_integration_tests/test_scenarios/cpp:cpp_test_scenarios -- --list-scenarios ``` diff --git a/feature_integration_tests/configs/BUILD b/feature_integration_tests/configs/BUILD index dce9a78284e..1bdd56eca30 100644 --- a/feature_integration_tests/configs/BUILD +++ b/feature_integration_tests/configs/BUILD @@ -15,6 +15,8 @@ exports_files( "dlt_config_qnx_x86_64.json", "dlt_config_x86_64.json", "qemu_bridge_config.json", + "lifecycle_daemon_config.json", + "lifecycle_daemon_retry_config.json", ], ) diff --git a/feature_integration_tests/configs/lifecycle_daemon_config.json b/feature_integration_tests/configs/lifecycle_daemon_config.json new file mode 100644 index 00000000000..5bb3e10f811 --- /dev/null +++ b/feature_integration_tests/configs/lifecycle_daemon_config.json @@ -0,0 +1,115 @@ +{ + "schema_version": 1, + "defaults": { + "deployment_config": { + "bin_dir": "__FIT_RUNTIME_ROOT__/bin", + "ready_timeout": 2.0, + "shutdown_timeout": 2.0, + "ready_recovery_action": { + "restart": { + "number_of_attempts": 2 + } + }, + "recovery_action": { + "switch_run_target": { + "run_target": "fallback_run_target" + } + }, + "sandbox": { + "uid": 1001, + "gid": 1001, + "scheduling_policy": "SCHED_OTHER", + "scheduling_priority": 0 + } + }, + "component_properties": { + "application_profile": { + "application_type": "Reporting", + "is_self_terminating": false, + "alive_supervision": { + "reporting_cycle": 0.1, + "min_indications": 1, + "max_indications": 3, + "failed_cycles_tolerance": 1 + } + }, + "ready_condition": { + "process_state": "Running" + } + } + }, + "components": { + "cpp_supervised_app": { + "component_properties": { + "binary_name": "cpp_supervised_app", + "application_profile": { + "application_type": "Reporting_And_Supervised" + }, + "process_arguments": [ + "-d50" + ] + }, + "deployment_config": { + "environmental_variables": { + "PROCESSIDENTIFIER": "cpp_supervised_app", + "IDENTIFIER": "cpp_supervised_app" + }, + "sandbox": { + "uid": 1001, + "gid": 1001, + "scheduling_policy": "SCHED_FIFO", + "scheduling_priority": 20 + } + } + }, + "rust_supervised_app": { + "component_properties": { + "binary_name": "rust_supervised_app", + "depends_on": [ + "cpp_supervised_app" + ], + "application_profile": { + "application_type": "Reporting_And_Supervised" + }, + "process_arguments": [ + "-d50" + ] + }, + "deployment_config": { + "environmental_variables": { + "PROCESSIDENTIFIER": "rust_supervised_app", + "IDENTIFIER": "rust_supervised_app" + }, + "sandbox": { + "uid": 1001, + "gid": 1001, + "scheduling_policy": "SCHED_RR", + "scheduling_priority": 10 + } + } + } + }, + "run_targets": { + "Startup": { + "depends_on": [ + "cpp_supervised_app", + "rust_supervised_app" + ], + "recovery_action": { + "switch_run_target": { + "run_target": "fallback_run_target" + } + } + } + }, + "initial_run_target": "Startup", + "alive_supervision": { + "evaluation_cycle": 0.05 + }, + "fallback_run_target": { + "depends_on": [ + "cpp_supervised_app", + "rust_supervised_app" + ] + } +} diff --git a/feature_integration_tests/configs/lifecycle_daemon_retry_config.json b/feature_integration_tests/configs/lifecycle_daemon_retry_config.json new file mode 100644 index 00000000000..c0e41057f66 --- /dev/null +++ b/feature_integration_tests/configs/lifecycle_daemon_retry_config.json @@ -0,0 +1,81 @@ +{ + "schema_version": 1, + "defaults": { + "deployment_config": { + "bin_dir": "__FIT_RUNTIME_ROOT__/bin", + "ready_timeout": 2.0, + "shutdown_timeout": 2.0, + "ready_recovery_action": { + "restart": { + "number_of_attempts": 0 + } + }, + "recovery_action": { + "switch_run_target": { + "run_target": "fallback_run_target" + } + }, + "sandbox": { + "uid": 1001, + "gid": 1001, + "scheduling_policy": "SCHED_OTHER", + "scheduling_priority": 0 + } + }, + "component_properties": { + "application_profile": { + "application_type": "Reporting", + "is_self_terminating": false, + "alive_supervision": { + "reporting_cycle": 0.1, + "min_indications": 1, + "max_indications": 3, + "failed_cycles_tolerance": 1 + } + }, + "ready_condition": { + "process_state": "Running" + } + } + }, + "components": { + "flaky_startup_app": { + "component_properties": { + "binary_name": "flaky_startup_app", + "process_arguments": [ + "__FIT_RUNTIME_ROOT__/flaky_startup_app.counter", + "__FIT_CRASHES_BEFORE_SUCCESS__" + ] + }, + "deployment_config": { + "ready_recovery_action": { + "restart": { + "number_of_attempts": 2 + } + }, + "environmental_variables": { + "PROCESSIDENTIFIER": "flaky_startup_app" + } + } + } + }, + "run_targets": { + "Startup": { + "depends_on": [ + "flaky_startup_app" + ], + "recovery_action": { + "switch_run_target": { + "run_target": "fallback_run_target" + } + } + } + }, + "initial_run_target": "Startup", + "alive_supervision": { + "evaluation_cycle": 0.05 + }, + "fallback_run_target": { + "depends_on": [] + } +} diff --git a/feature_integration_tests/test_cases/BUILD b/feature_integration_tests/test_cases/BUILD index 18d0b24212f..36bbcab4b09 100644 --- a/feature_integration_tests/test_cases/BUILD +++ b/feature_integration_tests/test_cases/BUILD @@ -37,9 +37,11 @@ compile_pip_requirements( ) # Tests targets + score_py_pytest( - name = "fit_rust", - srcs = glob(["tests/**/*.py"]), + name = "fit_rust_persistency", + timeout = "long", + srcs = glob(["tests/persistency/**/*.py"]), args = [ "-m rust", "--traces=all", @@ -60,8 +62,40 @@ score_py_pytest( ) score_py_pytest( - name = "fit_cpp", - srcs = glob(["tests/**/*.py"]), + name = "fit_rust_scenario_lifecycle", + timeout = "long", + srcs = ["tests/lifecycle/test_conditional_launching_scenario.py"], + args = [ + "-m rust", + "--traces=all", + "--rust-target-path=$(rootpath //feature_integration_tests/test_scenarios/rust:rust_test_scenarios)", + ], + data = [ + "conftest.py", + "fit_scenario.py", + "lifecycle_scenario.py", + "test_properties.py", + "//feature_integration_tests/test_scenarios/rust:rust_test_scenarios", + ], + env = { + "RUST_BACKTRACE": "1", + }, + pytest_config = "//:pyproject.toml", + deps = all_requirements, +) + +test_suite( + name = "fit_rust", + tests = [ + ":fit_rust_persistency", + ":fit_rust_scenario_lifecycle", + ], +) + +score_py_pytest( + name = "fit_cpp_persistency", + timeout = "long", + srcs = glob(["tests/persistency/**/*.py"]), args = [ "-m cpp", "--traces=all", @@ -78,10 +112,129 @@ score_py_pytest( deps = all_requirements, ) +score_py_pytest( + name = "fit_cpp_scenario_lifecycle", + timeout = "long", + srcs = ["tests/lifecycle/test_conditional_launching_scenario.py"], + args = [ + "-m cpp", + "--traces=all", + "--cpp-target-path=$(rootpath //feature_integration_tests/test_scenarios/cpp:cpp_test_scenarios)", + ], + data = [ + "conftest.py", + "fit_scenario.py", + "lifecycle_scenario.py", + "test_properties.py", + "//feature_integration_tests/test_scenarios/cpp:cpp_test_scenarios", + ], + pytest_config = "//:pyproject.toml", + deps = all_requirements, +) + +test_suite( + name = "fit_cpp", + tests = [ + ":fit_cpp_persistency", + ":fit_cpp_scenario_lifecycle", + ], +) + +# Daemon-driven lifecycle tests against a real launch_manager. No -m rust/cpp filter: these tests +# don't use the scenario binaries, and some carry no language marker, so a filter would deselect them. +score_py_pytest( + name = "fit_lifecycle_daemon", + timeout = "long", + srcs = [ + "tests/lifecycle/test_conditional_launching.py", + "tests/lifecycle/test_process_launching_with_daemon.py", + ], + args = [ + "--traces=all", + ], + data = [ + "conftest.py", + "daemon_helpers.py", + "test_properties.py", + "tests/lifecycle/conftest.py", + "//feature_integration_tests/configs:lifecycle_daemon_config.json", + "@flatbuffers//:flatc", + "@score_lifecycle//examples/cpp_supervised_app", + "@score_lifecycle//examples/rust_supervised_app", + "@score_lifecycle//score/launch_manager", + "@score_lifecycle//score/launch_manager/src/daemon/src/configuration:lm_flatcfg_fbs", + "@score_lifecycle//score/launch_manager/src/daemon/src/configuration/config_schema:launch_manager.schema.json", + "@score_lifecycle//scripts/config_mapping:lifecycle_config", + ], + env = { + "FIT_CPP_SUPERVISED_APP_PATH": "$(rootpath @score_lifecycle//examples/cpp_supervised_app)", + "FIT_LAUNCH_MANAGER_PATH": "$(rootpath @score_lifecycle//score/launch_manager)", + "FIT_FLATC_PATH": "$(rootpath @flatbuffers//:flatc)", + "FIT_LIFECYCLE_CONFIG_SCHEMA_PATH": "$(rootpath @score_lifecycle//score/launch_manager/src/daemon/src/configuration/config_schema:launch_manager.schema.json)", + "FIT_LIFECYCLE_CONFIG_TOOL_PATH": "$(rootpath @score_lifecycle//scripts/config_mapping:lifecycle_config)", + "FIT_LIFECYCLE_DAEMON_CONFIG_PATH": "$(rootpath //feature_integration_tests/configs:lifecycle_daemon_config.json)", + "FIT_LIFECYCLE_LM_SCHEMA_PATH": "$(rootpath @score_lifecycle//score/launch_manager/src/daemon/src/configuration:lm_flatcfg_fbs)", + "FIT_RUST_SUPERVISED_APP_PATH": "$(rootpath @score_lifecycle//examples/rust_supervised_app)", + "RUST_BACKTRACE": "1", + }, + env_inherit = ["FIT_ENABLE_SETCAP"], + pytest_config = "//:pyproject.toml", + # Under the default linux-sandbox, PR_SET_NO_NEW_PRIVS makes the setcap grant inert, so the + # uid/gid/scheduling tests skip (as in CI). They run only with --spawn_strategy=local and + # --test_env=FIT_ENABLE_SETCAP=1 (see README.md). + # "exclusive": launch_manager uses fixed POSIX shm names on the host-wide /dev/shm, so no other + # launch_manager-driven target may run concurrently. + tags = ["exclusive"], + deps = all_requirements, +) + +# Dedicated, isolated coverage for ready_recovery_action.restart.number_of_attempts +# (feat_req__lifecycle__retries_configurable). Each daemon invocation generates its own +# single-component config and runtime tree under TEST_TMPDIR. +score_py_pytest( + name = "fit_lifecycle_retries", + timeout = "long", + srcs = [ + "tests/lifecycle/test_retry_exhaustion.py", + ], + args = [ + "--traces=all", + ], + data = [ + "conftest.py", + "daemon_helpers.py", + "fit_scenario.py", + "lifecycle_scenario.py", + "test_properties.py", + "//feature_integration_tests/configs:lifecycle_daemon_retry_config.json", + "//feature_integration_tests/test_cases/support_apps/flaky_startup_app", + "@flatbuffers//:flatc", + "@score_lifecycle//score/launch_manager", + "@score_lifecycle//score/launch_manager/src/daemon/src/configuration:lm_flatcfg_fbs", + "@score_lifecycle//score/launch_manager/src/daemon/src/configuration/config_schema:launch_manager.schema.json", + "@score_lifecycle//scripts/config_mapping:lifecycle_config", + ], + env = { + "FIT_FLAKY_STARTUP_APP_PATH": "$(rootpath //feature_integration_tests/test_cases/support_apps/flaky_startup_app)", + "FIT_FLATC_PATH": "$(rootpath @flatbuffers//:flatc)", + "FIT_LAUNCH_MANAGER_PATH": "$(rootpath @score_lifecycle//score/launch_manager)", + "FIT_LIFECYCLE_CONFIG_SCHEMA_PATH": "$(rootpath @score_lifecycle//score/launch_manager/src/daemon/src/configuration/config_schema:launch_manager.schema.json)", + "FIT_LIFECYCLE_CONFIG_TOOL_PATH": "$(rootpath @score_lifecycle//scripts/config_mapping:lifecycle_config)", + "FIT_LIFECYCLE_LM_SCHEMA_PATH": "$(rootpath @score_lifecycle//score/launch_manager/src/daemon/src/configuration:lm_flatcfg_fbs)", + "FIT_LIFECYCLE_RETRY_CONFIG_PATH": "$(rootpath //feature_integration_tests/configs:lifecycle_daemon_retry_config.json)", + }, + pytest_config = "//:pyproject.toml", + # See fit_lifecycle_daemon: launch_manager's fixed /dev/shm names forbid concurrent targets. + tags = ["exclusive"], + deps = all_requirements, +) + test_suite( name = "fit", tests = [ ":fit_cpp", + ":fit_lifecycle_daemon", + ":fit_lifecycle_retries", ":fit_rust", ], ) diff --git a/feature_integration_tests/test_cases/conftest.py b/feature_integration_tests/test_cases/conftest.py index 662b7210943..36713ae4219 100644 --- a/feature_integration_tests/test_cases/conftest.py +++ b/feature_integration_tests/test_cases/conftest.py @@ -15,6 +15,39 @@ import pytest from testing_utils import BazelTools +try: + # Private API - not guaranteed stable across pytest versions. + from _pytest.mark.expression import Expression +except ImportError: + Expression = None + +_DEFAULT_RUST_TARGET = "//feature_integration_tests/test_scenarios/rust:rust_test_scenarios" + + +def _selected_versions(session: pytest.Session) -> set[str]: + """Return the scenario variants explicitly requested by the mark expression. + + Uses pytest's own marker expression evaluator so that logical operators and + negations are respected. For example, ``-m "not rust"`` must *not* select the + Rust build, while a plain substring check would incorrectly match it. + Falls back to all variants when no expression is given or parsing fails. + """ + mark_expression = session.config.option.markexpr or "" + if not mark_expression or Expression is None: + return {"rust", "cpp"} + try: + expr = Expression.compile(mark_expression) + # pytest 9's MatcherNameAdapter.__call__ forwards **kwargs to the matcher (needed for + # registered markers with args, e.g. `test_properties(x=1)`), so a matcher that only + # accepts `name` raises TypeError for an expression like `"rust and test_properties(x=1)"`. + # Keep evaluate() inside this try so that also falls back to all variants. + selected_versions = { + version for version in ("rust", "cpp") if expr.evaluate(lambda name, **_kwargs: name == version) + } + except Exception: # noqa: BLE001 – malformed expression; fall back to all variants + return {"rust", "cpp"} + return selected_versions or {"rust", "cpp"} + # Cmdline options def pytest_addoption(parser): @@ -31,7 +64,7 @@ def pytest_addoption(parser): parser.addoption( "--rust-target-name", type=str, - default="//feature_integration_tests/test_scenarios/rust:rust_test_scenarios", + default=_DEFAULT_RUST_TARGET, help="Rust test scenario executable target.", ) parser.addoption( @@ -88,18 +121,21 @@ def pytest_sessionstart(session): # Build scenarios. if session.config.getoption("--build-scenarios"): build_timeout = session.config.getoption("--build-scenarios-timeout") + selected_versions = _selected_versions(session) # Build Rust test scenarios. - print("Building Rust test scenarios executable...") - rust_tools = BazelTools(option_prefix="rust", build_timeout=build_timeout) - rust_target_name = session.config.getoption("--rust-target-name") - rust_tools.build(rust_target_name) + if "rust" in selected_versions: + print("Building Rust test scenarios executable...") + rust_tools = BazelTools(option_prefix="rust", build_timeout=build_timeout) + rust_target_name = session.config.getoption("--rust-target-name") + rust_tools.build(rust_target_name) # Build C++ test scenarios. - print("Building C++ test scenarios executable...") - cpp_tools = BazelTools(option_prefix="cpp", build_timeout=build_timeout) - cpp_target_name = session.config.getoption("--cpp-target-name") - cpp_tools.build(cpp_target_name) + if "cpp" in selected_versions: + print("Building C++ test scenarios executable...") + cpp_tools = BazelTools(option_prefix="cpp", build_timeout=build_timeout) + cpp_target_name = session.config.getoption("--cpp-target-name") + cpp_tools.build(cpp_target_name) except Exception as e: pytest.exit(str(e), returncode=1) diff --git a/feature_integration_tests/test_cases/daemon_helpers.py b/feature_integration_tests/test_cases/daemon_helpers.py new file mode 100644 index 00000000000..0d0ae45257a --- /dev/null +++ b/feature_integration_tests/test_cases/daemon_helpers.py @@ -0,0 +1,797 @@ +# ******************************************************************************* +# Copyright (c) 2026 Contributors to the Eclipse Foundation +# +# See the NOTICE file(s) distributed with this work for additional +# information regarding copyright ownership. +# +# This program and the accompanying materials are made available under the +# terms of the Apache License Version 2.0 which is available at +# https://www.apache.org/licenses/LICENSE-2.0 +# +# SPDX-License-Identifier: Apache-2.0 +# ******************************************************************************* +"""Daemon helpers for lifecycle behavior tests against real Launch Manager.""" + +from __future__ import annotations + +import json +import os +import re +import shutil +import signal +import subprocess +import tempfile +import threading +import time +from dataclasses import dataclass +from pathlib import Path +from typing import Any + +import pytest + +_TARGET_ENV_MAP = { + "@score_lifecycle//score/launch_manager:launch_manager": "FIT_LAUNCH_MANAGER_PATH", + "@score_lifecycle//examples/rust_supervised_app:rust_supervised_app": "FIT_RUST_SUPERVISED_APP_PATH", + "@score_lifecycle//examples/cpp_supervised_app:cpp_supervised_app": "FIT_CPP_SUPERVISED_APP_PATH", + "//feature_integration_tests/configs:lifecycle_daemon_config.json": "FIT_LIFECYCLE_DAEMON_CONFIG_PATH", + "//feature_integration_tests/test_cases/support_apps/flaky_startup_app:flaky_startup_app": ( + "FIT_FLAKY_STARTUP_APP_PATH" + ), + "//feature_integration_tests/configs:lifecycle_daemon_retry_config.json": "FIT_LIFECYCLE_RETRY_CONFIG_PATH", + "@score_lifecycle//scripts/config_mapping:lifecycle_config": "FIT_LIFECYCLE_CONFIG_TOOL_PATH", + "@score_lifecycle//score/launch_manager/src/daemon/src/configuration/config_schema:launch_manager.schema.json": "FIT_LIFECYCLE_CONFIG_SCHEMA_PATH", + "@score_lifecycle//score/launch_manager/src/daemon/src/configuration:lm_flatcfg_fbs": "FIT_LIFECYCLE_LM_SCHEMA_PATH", + "@flatbuffers//:flatc": "FIT_FLATC_PATH", +} + + +def _resolve_from_env(target: str) -> Path | None: + """Resolve a target path from Bazel-provided runfile environment variables.""" + env_var = _TARGET_ENV_MAP.get(target) + if env_var is None: + return None + + raw_path = os.environ.get(env_var) + if not raw_path: + return None + + candidate = Path(raw_path) + search_roots = [Path.cwd()] + + test_srcdir = os.environ.get("TEST_SRCDIR") + test_workspace = os.environ.get("TEST_WORKSPACE") + if test_srcdir and test_workspace: + search_roots.append(Path(test_srcdir) / test_workspace) + if test_srcdir: + search_roots.append(Path(test_srcdir)) + + for root in search_roots: + resolved = candidate if candidate.is_absolute() else (root / candidate) + if resolved.exists(): + return resolved.resolve() + + return None + + +def _resolve_target_path(target: str) -> Path: + """Resolve an executable/file path from a bazel target label via its runfile env var.""" + env_resolved = _resolve_from_env(target) + if env_resolved is not None: + return env_resolved + + env_var = _TARGET_ENV_MAP.get(target) + raise RuntimeError( + f"Could not resolve target {target!r}: environment variable " + f"{env_var!r} is not set or does not point to an existing file. " + "Ensure the corresponding data dependency is declared on the test target." + ) + + +def pgrep_cmdline_pattern(binary_path: str) -> str: + """Build POSIX ERE pattern matching binary with optional arguments.""" + return rf"^{re.escape(binary_path)}([[:space:]]|$)" + + +def is_running(binary_path: str | Path) -> bool: + """True if some process's cmdline starts with `binary_path` (pgrep). This is process + existence only, not launch_manager's reported Running state.""" + result = subprocess.run( + ["pgrep", "-f", pgrep_cmdline_pattern(str(binary_path))], + capture_output=True, + text=True, + check=False, + ) + return result.returncode == 0 + + +def first_pid(binary_path: str | Path) -> str | None: + """First pgrep match for `binary_path` (see `is_running`), or None.""" + result = subprocess.run( + ["pgrep", "-f", pgrep_cmdline_pattern(str(binary_path))], + capture_output=True, + text=True, + check=False, + ) + if result.returncode != 0: + return None + lines = [line for line in result.stdout.splitlines() if line] + return lines[0] if lines else None + + +def wait_until(predicate, timeout_s: float, interval_s: float = 0.2) -> bool: + """Poll `predicate` until it is truthy (True) or `timeout_s` elapses (False).""" + deadline = time.time() + timeout_s + while time.time() < deadline: + if predicate(): + return True + time.sleep(interval_s) + return False + + +# cap_setuid/cap_setgid: apply sandbox uid/gid; cap_sys_nice: apply non-SCHED_OTHER policies; +# cap_kill: launch_manager (runner uid) must signal children running as `_SANDBOX_UID`. +_SETCAP_CAPS = "cap_setuid,cap_setgid,cap_sys_nice,cap_kill+ep" + +# Sandbox uid substituted when capabilities are granted. Must differ from the runner's uid, else +# the uid test cannot tell an applied identity from an inherited one (config's 1001 is commonly +# the runner's own uid). The gid is NOT remapped to a distinct value: see `_generate_runtime_config`. +_SANDBOX_UID = 65533 + +# Capability-granted copies of `kill` (cap_kill) and `cat` (cap_sys_ptrace, cap_dac_read_search), +# staged per daemon by `_spawn_daemon` and deleted by `_teardown`. They let the runner signal apps +# running as `_SANDBOX_UID` and read their /proc//environ with only the setcap sudoers rule. +_privileged_kill: Path | None = None +_privileged_cat: Path | None = None + + +def _mount_nosuid(path: Path) -> bool: + """Best-effort check (via `findmnt`) whether `path` is on a `nosuid` mount, which drops + file capabilities at exec time. Used only to enrich a failed-grant diagnostic.""" + try: + findmnt = shutil.which("findmnt") + if findmnt is None: + return False + result = subprocess.run( + [findmnt, "-n", "-o", "OPTIONS", "-T", str(path)], + capture_output=True, + text=True, + check=False, + ) + return result.returncode == 0 and "nosuid" in result.stdout + except OSError: + return False + + +def _grant_sandbox_capabilities( + binary_path: Path, + caps: str = _SETCAP_CAPS, + required: tuple[str, ...] = ("cap_setuid", "cap_setgid"), +) -> tuple[bool, str]: + """Best-effort `setcap caps binary_path`. Never raises. + + Returns `(granted, reason)`. `granted` is True only if `getcap` reads back every cap in + `required` (when `getcap` is available); `reason` is a diagnostic suitable for a skip message. + + Tries `sudo -n setcap` first when FIT_ENABLE_SETCAP=1 (needs a passwordless sudoers rule for + the setcap binary with no pinned arguments), then plain `setcap` (succeeds only as root). + Under `bazel test` the variable must be passed with `--test_env`, and the grant is inert + inside linux-sandbox (PR_SET_NO_NEW_PRIVS) even when setcap succeeds. + + Limitation: only `required` is verified; the other caps in `caps` are assumed granted with it. + """ + if shutil.which("setcap") is None: + return False, "setcap binary not found on PATH" + + setcap_enabled = os.environ.get("FIT_ENABLE_SETCAP") == "1" + attempts: list[tuple[list[str], str]] = [ + (["setcap", caps, str(binary_path)], "plain setcap (requires running as root)") + ] + if setcap_enabled: + if shutil.which("sudo") is None: + attempts.append(([], "FIT_ENABLE_SETCAP=1 set but 'sudo' not found on PATH")) + else: + attempts.insert( + 0, + (["sudo", "-n", "setcap", caps, str(binary_path)], "sudo -n setcap"), + ) + else: + attempts.append(([], "FIT_ENABLE_SETCAP not set to '1'; skipping sudo setcap attempt")) + + failures: list[str] = [] + for cmd, label in attempts: + if not cmd: + failures.append(label) + continue + result = subprocess.run(cmd, capture_output=True, text=True, check=False) + if result.returncode != 0: + failures.append( + f"{label} failed (rc={result.returncode}): " + f"{result.stderr.strip() or result.stdout.strip() or ''}" + ) + continue + + # setcap can report success while the kernel still drops the capability at exec + # time (e.g. the binary lives on a filesystem mounted `nosuid`). Verify by reading + # the xattr back instead of trusting the exit code. + getcap = shutil.which("getcap") + if getcap is not None: + verify = subprocess.run([getcap, str(binary_path)], capture_output=True, text=True, check=False) + if any(cap not in verify.stdout for cap in required): + nosuid_hint = " (path is on a 'nosuid' mount)" if _mount_nosuid(binary_path) else "" + failures.append( + f"{label} reported success but getcap did not confirm the capabilities" + f"{nosuid_hint}: {verify.stdout.strip() or ''}" + ) + continue + + return True, f"granted via {label}" + + return False, "; ".join(failures) if failures else "no grant attempt produced a result" + + +def signal_process(pid: str, sig: str, *, sandbox_privileged: bool) -> tuple[bool, str]: + """Send `sig` (e.g. "-9", "-STOP", "-CONT") to `pid`. Returns `(sent, reason)`; never raises. + + Tries plain `kill`, then (if `sandbox_privileged`) the staged cap_kill copy, then + `sudo -n kill` when FIT_ENABLE_SETCAP=1. The sudo fallback only works with a sudoers rule + for `kill`, which the documented setup does not provide. + """ + attempts: list[list[str]] = [["kill", sig, pid]] + if sandbox_privileged and _privileged_kill is not None: + attempts.append([str(_privileged_kill), sig, pid]) + if sandbox_privileged and os.environ.get("FIT_ENABLE_SETCAP") == "1" and shutil.which("sudo") is not None: + attempts.append(["sudo", "-n", "kill", sig, pid]) + + failures: list[str] = [] + for cmd in attempts: + result = subprocess.run(cmd, capture_output=True, text=True, check=False) + if result.returncode == 0: + return True, f"sent via {' '.join(cmd)}" + failures.append(f"{' '.join(cmd)} failed (rc={result.returncode}): {result.stderr.strip() or ''}") + + return False, "; ".join(failures) + + +def _wait_for_apps(apps: dict[str, Path], timeout_s: float = 8.0, interval_s: float = 0.2) -> bool: + return wait_until(lambda: all(is_running(path) for path in apps.values()), timeout_s, interval_s) + + +def _tmpdir_root() -> Path: + """Return the writable temp root for the current test invocation. + + Bazel sets `TEST_TMPDIR` to a fresh directory per test. Outside Bazel, fall + back to the system temp dir; callers create unique children in either case. + """ + value = os.environ.get("TEST_TMPDIR") + if value: + return Path(value) + return Path(tempfile.gettempdir()) + + +@dataclass +class ManagedDaemon: + """A launch_manager subprocess plus the stdout/stderr lines its reader thread collected.""" + + process: subprocess.Popen[str] + _lines: list[str] + _thread: threading.Thread + + def is_running(self) -> bool: + return self.process.poll() is None + + def pid(self) -> int: + return self.process.pid + + def stop(self) -> None: + """SIGTERM the daemon's process group, SIGKILL after 5 s. Does not reach supervised + apps, which launch_manager moves into their own process groups (see `_teardown`).""" + if self.is_running(): + os.killpg(os.getpgid(self.process.pid), signal.SIGTERM) + deadline = time.time() + 5.0 + while self.is_running() and time.time() < deadline: + time.sleep(0.1) + if self.is_running(): + os.killpg(os.getpgid(self.process.pid), signal.SIGKILL) + self.process.wait(timeout=5) + + def close_output(self) -> None: + """Join the stdout reader (1 s) and close the pipe. + + Call only after the supervised apps are dead: they inherit the pipe, so the reader sees + EOF only then. If the reader is still alive the pipe is left open (leaked) rather than + closed under it, which would deadlock. + """ + self._thread.join(timeout=1) + if self.process.stdout is not None and not self._thread.is_alive(): + self.process.stdout.close() + + def get_logs(self) -> str: + return "\n".join(self._lines) + + +def _cleanup_runtime_root(runtime_root: Path) -> None: + """Remove a daemon's uniquely allocated runtime directory.""" + shutil.rmtree(runtime_root, ignore_errors=True) + + +def _generate_runtime_config( + config_template: str, + runtime_root: Path, + etc_dir: Path, + sandbox_privileged: bool = True, + *, + remap_sandbox_uid: bool = False, + independent_apps: bool = False, + ready_timeout_s: float | None = None, + crashes_before_success: int | None = None, +) -> None: + """Render `config_template` for one daemon and serialize it to `etc_dir`. + + Writes the rendered JSON to `etc_dir/lifecycle_config.json` (tests read expected values from + it) and the flatbuffer to `etc_dir/launch_manager_config.bin` via the upstream config mapper + and `flatc`. Raises RuntimeError with the tool's output if either step fails. + + Rendering always sets `bin_dir` to `runtime_root/bin` and the flaky-app counter path; optional + variants (rendered, not kept as copied config files, so they cannot drift): + - `independent_apps`: drop every component's `depends_on`. + - `ready_timeout_s`: override `defaults.deployment_config.ready_timeout`. + - `crashes_before_success`: fill the `__FIT_CRASHES_BEFORE_SUCCESS__` argument. + - `remap_sandbox_uid`: set every sandbox uid to `_SANDBOX_UID` and gid to the runner's gid, + and chmod `runtime_root` 0750. Limitation: the gid then equals the runner's, so a gid + check cannot prove setgid() was applied. + - `sandbox_privileged=False`: downgrade non-`SCHED_OTHER` policies to `SCHED_OTHER`/0. + launch_manager treats a failed sched_setscheduler() as fatal for the component, so without + CAP_SYS_NICE every such component would crash-loop. + """ + config = json.loads(_resolve_target_path(config_template).read_text(encoding="utf-8")) + config["defaults"]["deployment_config"]["bin_dir"] = str(runtime_root / "bin") + if ready_timeout_s is not None: + config["defaults"]["deployment_config"]["ready_timeout"] = ready_timeout_s + if independent_apps: + for component in config["components"].values(): + component["component_properties"].pop("depends_on", None) + + sandboxes = [config["defaults"]["deployment_config"].get("sandbox")] + sandboxes += [component.get("deployment_config", {}).get("sandbox") for component in config["components"].values()] + if remap_sandbox_uid: + for sandbox in filter(None, sandboxes): + sandbox["uid"] = _SANDBOX_UID + sandbox["gid"] = os.getgid() + runtime_root.chmod(0o750) + if not sandbox_privileged: + for sandbox in sandboxes: + if sandbox and sandbox.get("scheduling_policy") not in (None, "SCHED_OTHER"): + sandbox["scheduling_policy"] = "SCHED_OTHER" + sandbox["scheduling_priority"] = 0 + + placeholders = {"__FIT_RUNTIME_ROOT__/flaky_startup_app.counter": str(runtime_root / "flaky_startup_app.counter")} + if crashes_before_success is not None: + placeholders["__FIT_CRASHES_BEFORE_SUCCESS__"] = str(crashes_before_success) + for component in config["components"].values(): + arguments = component["component_properties"].get("process_arguments", []) + component["component_properties"]["process_arguments"] = [placeholders.get(a, a) for a in arguments] + + rendered_config = etc_dir / "lifecycle_config.json" + rendered_config.write_text(json.dumps(config), encoding="utf-8") + generated_dir = etc_dir / "generated" + generated_dir.mkdir() + config_tool = _resolve_target_path("@score_lifecycle//scripts/config_mapping:lifecycle_config") + config_schema = _resolve_target_path( + "@score_lifecycle//score/launch_manager/src/daemon/src/configuration/config_schema:launch_manager.schema.json" + ) + config_mapping_result = subprocess.run( + [str(config_tool), str(rendered_config), "--schema", str(config_schema), "-o", str(generated_dir)], + capture_output=True, + text=True, + check=False, + ) + if config_mapping_result.returncode != 0: + raise RuntimeError( + f"Command failed (rc={config_mapping_result.returncode}): {config_tool} {rendered_config} " + f"--schema {config_schema} -o {generated_dir}\n" + f"stdout:\n{config_mapping_result.stdout}\nstderr:\n{config_mapping_result.stderr}" + ) + + flatc = _resolve_target_path("@flatbuffers//:flatc") + lm_schema = _resolve_target_path( + "@score_lifecycle//score/launch_manager/src/daemon/src/configuration:lm_flatcfg_fbs" + ) + generated_config = generated_dir / f"{rendered_config.stem}_gen.json" + # launch_manager defaults to loading "etc/launch_manager_config.bin", and flatc names its + # output after the input file's stem, so the input must be named to match. + flatc_input = generated_dir / "launch_manager_config.json" + shutil.copy2(generated_config, flatc_input) + flatc_cmd = [ + str(flatc), + "--binary", + "--strict-json", + "-o", + str(etc_dir), + str(lm_schema), + str(flatc_input), + ] + flatc_result = subprocess.run( + flatc_cmd, + capture_output=True, + text=True, + check=False, + ) + if flatc_result.returncode != 0: + raise RuntimeError( + f"Command failed (rc={flatc_result.returncode}): {' '.join(flatc_cmd)}\n" + f"stdout:\n{flatc_result.stdout}\nstderr:\n{flatc_result.stderr}" + ) + + +# launch_manager creates POSIX shm objects with fixed names ("/ipc_shared_mem", "/_nudge~._.~me_") +# using O_CREAT|O_EXCL on the host-wide /dev/shm (not isolated by `unshare -i`). A second daemon +# started before the first has unlinked them gets EEXIST and silently launches nothing. So daemon +# lifetimes must not overlap: this registry enforces it within one pytest process; across Bazel +# test processes the lifecycle targets are tagged "exclusive". +_live_daemons: list[ManagedDaemon] = [] + + +def _assert_no_live_daemon() -> None: + """Fail the test if a daemon started by this process is still running.""" + _live_daemons[:] = [d for d in _live_daemons if d.is_running()] + if _live_daemons: + pytest.fail( + f"Another launch_manager (pid={_live_daemons[0].pid()}) is still running; overlapping " + "daemon lifetimes collide on launch_manager's fixed POSIX shm names. Don't start a " + "daemon from a test that also holds the class-scoped `launch_manager_daemon` fixture." + ) + + +def _spawn_daemon( + work_dir: Path, + etc_dir: Path, + runtime_root: Path, + config_template: str, + staged_binaries: list[tuple[Path, Path, int]], + grant_sandbox_capabilities: bool = False, + **config_options: Any, +) -> tuple[ManagedDaemon, bool, str]: + """Copy launch_manager (0700) and `staged_binaries` (src, dst, mode) into place, render the + config (`config_options` go to `_generate_runtime_config`), and start launch_manager in its + own session with stdout+stderr collected by a reader thread. + + With `grant_sandbox_capabilities`, tries to setcap launch_manager; if that succeeds it also + stages the privileged `kill`/`cat` copies, and only when both are staged remaps the sandbox + uid. Returns `(daemon, sandbox_privileged, sandbox_privileged_reason)`. + + `pytest.fail`s if another daemon from this process is alive or if launch_manager exits within + the 1 s startup window; in the latter case (or any exception after Popen) it tears down the + daemon itself, since the caller never receives it. + """ + _assert_no_live_daemon() + launch_manager = _resolve_target_path("@score_lifecycle//score/launch_manager:launch_manager") + lm_dst = work_dir / "launch_manager" + shutil.copy2(launch_manager, lm_dst) + # 0700: once granted cap_setuid it must not be executable by other local users. + lm_dst.chmod(0o700) + + global _privileged_kill, _privileged_cat + _privileged_kill = _privileged_cat = None + if grant_sandbox_capabilities: + sandbox_privileged, sandbox_privileged_reason = _grant_sandbox_capabilities(lm_dst) + else: + sandbox_privileged, sandbox_privileged_reason = False, "not requested" + if sandbox_privileged: + # Remap the uid only if the runner can still signal the apps and read their environ. + kill_tool, kill_reason = _stage_privileged_tool(work_dir, "kill", ("cap_kill",)) + cat_tool, cat_reason = _stage_privileged_tool(work_dir, "cat", ("cap_sys_ptrace", "cap_dac_read_search")) + if kill_tool is not None and cat_tool is not None: + _privileged_kill, _privileged_cat = kill_tool, cat_tool + else: + for tool in (kill_tool, cat_tool): + if tool is not None: + tool.unlink() + sandbox_privileged_reason += f"; sandbox uid not remapped: {kill_reason}; {cat_reason}" + + for src, dst, mode in staged_binaries: + shutil.copy2(src, dst) + dst.chmod(mode) + + _generate_runtime_config( + config_template, + runtime_root, + etc_dir, + sandbox_privileged=sandbox_privileged, + remap_sandbox_uid=_privileged_kill is not None, + **config_options, + ) + + env = os.environ.copy() + env.setdefault("ECUCFG_ENV_VAR_ROOTFOLDER", str(etc_dir)) + + lines: list[str] = [] + process = subprocess.Popen( + [str(lm_dst)], + cwd=work_dir, + env=env, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + text=True, + start_new_session=True, + ) + + def _collect_output() -> None: + assert process.stdout is not None + for line in process.stdout: + line = line.rstrip("\n") + if line: + lines.append(line) + + thread = threading.Thread(target=_collect_output, daemon=True) + thread.start() + + daemon = ManagedDaemon(process=process, _lines=lines, _thread=thread) + _live_daemons.append(daemon) + + # The caller never receives `daemon` if this raises, so tear it down here; otherwise a + # detached launch_manager keeps running and blocks every later start via _live_daemons. + try: + # Startup window: a broken config makes launch_manager exit within it. + time.sleep(1.0) + if not daemon.is_running(): + pytest.fail(f"launch_manager failed to start. Logs:\n{daemon.get_logs()}") + except BaseException: + _teardown(daemon, [dst for _, dst, _ in staged_binaries], runtime_root) + raise + + return daemon, sandbox_privileged, sandbox_privileged_reason + + +def start_launch_manager_daemon( + tmp_path_factory: pytest.TempPathFactory, + blocked_apps: frozenset[str] = frozenset(), + stalled_apps: frozenset[str] = frozenset(), + wait_for_apps: bool = True, + independent_apps: bool = False, + ready_timeout_s: float | None = None, +) -> dict[str, Any]: + """Start launch_manager on lifecycle_daemon_config.json with rust_ and cpp_supervised_app. + + Always attempts the sandbox capability grant (see `_spawn_daemon`); check the returned + `sandbox_privileged` before asserting on uid/gid/scheduling. + - `blocked_apps` ("rust"/"cpp"): staged with mode 0000, so exec fails until the caller + chmods them 0755. Used to withhold a dependency. + - `stalled_apps`: replaced by a stub that runs but never reports Running, so launch_manager + waits out `ready_timeout` on it. + - `independent_apps`, `ready_timeout_s`: rendered into the config (`_generate_runtime_config`). + - `wait_for_apps`: `pytest.fail` unless every non-blocked app is running within 8 s. + Limitation: "running" means a matching process exists (pgrep), not launch_manager's Running state. + + Uses a fresh runtime root under `TEST_TMPDIR`; must not overlap another live daemon. + Returns the daemon info dict consumed by `stop_launch_manager_daemon`. + """ + + runtime_root = Path(tempfile.mkdtemp(prefix="lifecycle_fit-", dir=_tmpdir_root())) + bin_dir = runtime_root / "bin" + apps = { + "rust": bin_dir / "rust_supervised_app", + "cpp": bin_dir / "cpp_supervised_app", + } + daemon = None + try: + work_dir = tmp_path_factory.mktemp("lm-daemon") + etc_dir = work_dir / "etc" + etc_dir.mkdir(parents=True, exist_ok=True) + + bin_dir.mkdir(parents=True, exist_ok=True) + + rust_supervised = _resolve_target_path("@score_lifecycle//examples/rust_supervised_app:rust_supervised_app") + cpp_supervised = _resolve_target_path("@score_lifecycle//examples/cpp_supervised_app:cpp_supervised_app") + stall_stub = work_dir / "stall_stub.sh" + # `exec -a "$0"` keeps the staged app path as argv[0], so the anchored pgrep/pkill pattern + # matches the stub, and it stays a single process that dies on SIGTERM. + stall_stub.write_text('#!/bin/bash\nexec -a "$0" sleep infinity\n', encoding="utf-8") + staged_binaries = [ + ( + stall_stub if key in stalled_apps else src, + bin_dir / src.name, + 0o000 if key in blocked_apps else 0o755, + ) + for key, src in (("rust", rust_supervised), ("cpp", cpp_supervised)) + ] + + daemon, sandbox_privileged, sandbox_privileged_reason = _spawn_daemon( + work_dir, + etc_dir, + runtime_root, + "//feature_integration_tests/configs:lifecycle_daemon_config.json", + staged_binaries, + grant_sandbox_capabilities=True, + independent_apps=independent_apps, + ready_timeout_s=ready_timeout_s, + ) + + if wait_for_apps and not _wait_for_apps({k: v for k, v in apps.items() if k not in blocked_apps}): + process_snapshot = subprocess.run( + ["ps", "-eo", "pid,args"], + capture_output=True, + text=True, + check=False, + ) + pytest.fail( + "Launch Manager did not bring supervised apps to running state within timeout.\n" + f"Expected apps: {apps}\n" + f"Daemon logs:\n{daemon.get_logs()}\n" + f"Process snapshot (rc={process_snapshot.returncode}):\n" + f"{process_snapshot.stdout}{process_snapshot.stderr}" + ) + except BaseException: + # Kill apps too: any that already started outlive the daemon (own process group). + _teardown(daemon, list(apps.values()), runtime_root) + raise + + return { + "daemon": daemon, + "work_dir": work_dir, + "bin_dir": bin_dir, + "apps": apps, + "sandbox_privileged": sandbox_privileged, + "sandbox_privileged_reason": sandbox_privileged_reason, + "runtime_root": runtime_root, + "runtime_config": etc_dir / "lifecycle_config.json", + } + + +def start_flaky_retry_daemon( + tmp_path_factory: pytest.TempPathFactory, + crashes_before_success: int, +) -> dict[str, Any]: + """Start launch_manager on lifecycle_daemon_retry_config.json with only `flaky_startup_app`. + + The app aborts on its first `crashes_before_success` launches and then stays up, counting + every launch in `counter_path`, so tests can observe `number_of_attempts` deterministically. + No capability grant is requested. Does not wait for the app: whether it ever runs is what + the caller checks. Returns the daemon info dict consumed by `stop_flaky_retry_daemon`. + """ + runtime_root = Path(tempfile.mkdtemp(prefix="lifecycle_fit_retries-", dir=_tmpdir_root())) + bin_dir = runtime_root / "bin" + app_dst = bin_dir / "flaky_startup_app" + daemon = None + try: + work_dir = tmp_path_factory.mktemp("lm-retry-daemon") + etc_dir = work_dir / "etc" + etc_dir.mkdir(parents=True, exist_ok=True) + + bin_dir.mkdir(parents=True, exist_ok=True) + + flaky_app = _resolve_target_path( + "//feature_integration_tests/test_cases/support_apps/flaky_startup_app:flaky_startup_app" + ) + staged_binaries = [(flaky_app, app_dst, 0o755)] + + # runtime_root is a fresh mkdtemp, so the counter starts at 0. + counter_path = runtime_root / "flaky_startup_app.counter" + + daemon, _, _ = _spawn_daemon( + work_dir, + etc_dir, + runtime_root, + "//feature_integration_tests/configs:lifecycle_daemon_retry_config.json", + staged_binaries, + crashes_before_success=crashes_before_success, + ) + except BaseException: + _teardown(daemon, [app_dst], runtime_root) + raise + + return { + "daemon": daemon, + "work_dir": work_dir, + "bin_dir": bin_dir, + "app_path": app_dst, + "counter_path": counter_path, + "crashes_before_success": crashes_before_success, + "runtime_root": runtime_root, + "runtime_config": etc_dir / "lifecycle_config.json", + } + + +def _teardown(daemon: ManagedDaemon | None, app_paths: list[Path], runtime_root: Path) -> None: + """Stop `daemon` (if any), SIGKILL each of `app_paths` by cmdline, close the daemon's output, + delete the privileged tool copies and remove `runtime_root`. Each step runs even if an + earlier one raises. + + Apps are killed explicitly because they run in their own process groups, which `stop()` + does not reach. + """ + try: + if daemon is not None: + daemon.stop() + finally: + try: + for app_path in app_paths: + _kill_app(app_path) + if daemon is not None: + daemon.close_output() + finally: + _drop_privileged_tools() + _cleanup_runtime_root(runtime_root) + + +def _drop_privileged_tools() -> None: + """Delete the capability-granted `kill`/`cat` copies. `work_dir` is not removed (pytest keeps + recent basetemps), so they would otherwise stay on disk.""" + global _privileged_kill, _privileged_cat + for tool in (_privileged_kill, _privileged_cat): + if tool is not None: + tool.unlink(missing_ok=True) + _privileged_kill = _privileged_cat = None + + +def _stage_privileged_tool(work_dir: Path, name: str, caps: tuple[str, ...]) -> tuple[Path | None, str]: + """Copy the `name` binary from PATH to `work_dir/privileged/name` (mode 0700, runner only) and + grant it `caps` via `_grant_sandbox_capabilities`. Returns `(path, reason)`; `path` is None + if the grant was not verified (the copy is then left for the caller to delete). + + The basename is kept because procps `kill` dispatches on argv[0]. + """ + src = shutil.which(name) + if src is None: + return None, f"{name} binary not found on PATH" + dst = work_dir / "privileged" / name + dst.parent.mkdir(exist_ok=True) + shutil.copy2(Path(src).resolve(), dst) + dst.chmod(0o700) + granted, reason = _grant_sandbox_capabilities(dst, ",".join(caps) + "+ep", caps) + return (dst if granted else None), f"{'/'.join(caps)} on {name}: {reason}" + + +def read_proc_file(pid: str, name: str) -> bytes: + """Read `/proc//`, via the privileged `cat` copy when one is staged. + + Needed for files like `environ`: once the app runs as `_SANDBOX_UID` it is non-dumpable and + the runner lacks ptrace access to it. Raises RuntimeError (with stderr) if `cat` fails. + """ + path = f"/proc/{pid}/{name}" + if _privileged_cat is None: + return Path(path).read_bytes() + result = subprocess.run([str(_privileged_cat), path], capture_output=True, check=False) + if result.returncode != 0: + raise RuntimeError( + f"Command failed (rc={result.returncode}): {_privileged_cat} {path}\n" + f"stderr:\n{result.stderr.decode(errors='replace')}" + ) + return result.stdout + + +def _kill_app(app_path: Path) -> None: + """SIGKILL every process whose cmdline starts with `app_path`. Uses the cap_kill `kill` copy + when staged, since a plain pkill gets EPERM for apps running as `_SANDBOX_UID`. Best effort.""" + if _privileged_kill is None: + subprocess.run( + ["pkill", "-9", "-f", pgrep_cmdline_pattern(str(app_path))], capture_output=True, text=True, check=False + ) + return + pids = subprocess.run( + ["pgrep", "-f", pgrep_cmdline_pattern(str(app_path))], capture_output=True, text=True, check=False + ).stdout.split() + if pids: + subprocess.run([str(_privileged_kill), "-9", *pids], capture_output=True, text=True, check=False) + + +def _stop_daemon(daemon_info: dict[str, Any], app_paths: list[Path]) -> None: + """Tear down a started daemon, its `app_paths`, and its runtime root.""" + _teardown(daemon_info["daemon"], app_paths, daemon_info["runtime_root"]) + + +def stop_flaky_retry_daemon(daemon_info: dict[str, Any]) -> None: + """Tear down a daemon started by `start_flaky_retry_daemon`.""" + _stop_daemon(daemon_info, [daemon_info["app_path"]]) + + +def read_retry_attempt_count(counter_path: Path) -> int: + """Read flaky_startup_app's persisted attempt counter; 0 if it hasn't run yet.""" + try: + return int(counter_path.read_text().strip()) + except (FileNotFoundError, ValueError): + return 0 + + +def stop_launch_manager_daemon(daemon_info: dict[str, Any]) -> None: + """Tear down a daemon started by `start_launch_manager_daemon`.""" + _stop_daemon(daemon_info, list(daemon_info["apps"].values())) diff --git a/feature_integration_tests/test_cases/lifecycle_scenario.py b/feature_integration_tests/test_cases/lifecycle_scenario.py new file mode 100644 index 00000000000..69ec6cc4af0 --- /dev/null +++ b/feature_integration_tests/test_cases/lifecycle_scenario.py @@ -0,0 +1,71 @@ +# ******************************************************************************* +# Copyright (c) 2026 Contributors to the Eclipse Foundation +# +# See the NOTICE file(s) distributed with this work for additional +# information regarding copyright ownership. +# +# This program and the accompanying materials are made available under the +# terms of the Apache License Version 2.0 which is available at +# https://www.apache.org/licenses/LICENSE-2.0 +# +# SPDX-License-Identifier: Apache-2.0 +# ******************************************************************************* +""" +Base classes for lifecycle FITs. + +``LifecycleScenario``: ``FitScenario`` base for the scenario-binary tests +(test_conditional_launching_scenario.py); adds the class-scoped ``temp_dir`` fixture. + +``RetryDaemonScenario``: base for the flaky-retry daemon tests (test_retry_exhaustion.py), +which run no scenario binary; adds the class-scoped ``retry_daemon`` fixture. +""" + +from collections.abc import Generator +from pathlib import Path +from typing import Any + +import pytest +from fit_scenario import FitScenario, temp_dir_common + + +class LifecycleScenario(FitScenario): + """Base for lifecycle scenario-binary test classes; provides ``temp_dir``.""" + + @pytest.fixture(scope="class") + def temp_dir( + self, + tmp_path_factory: pytest.TempPathFactory, + version: str, + ) -> Generator[Path, None, None]: + """ + Per-class, per-version temporary directory for the scenario run. + + Parameters + ---------- + tmp_path_factory : pytest.TempPathFactory + Built-in pytest factory for temporary directories. + version : str + Parametrized scenario version (``"rust"`` or ``"cpp"``). + """ + yield from temp_dir_common(tmp_path_factory, self.__class__.__name__, version) + + +class RetryDaemonScenario: + """Base for flaky-retry daemon test classes. + + Subclasses set ``crashes_before_success``; ``retry_daemon`` starts one launch_manager per + class via ``daemon_helpers.start_flaky_retry_daemon`` and tears it down afterwards. + """ + + crashes_before_success: int + + @pytest.fixture(scope="class") + def retry_daemon(self, tmp_path_factory: pytest.TempPathFactory) -> Generator[dict[str, Any], None, None]: + # Lazy import: the scenario-lifecycle targets import this module without shipping daemon_helpers.py. + from daemon_helpers import start_flaky_retry_daemon, stop_flaky_retry_daemon + + daemon_info = start_flaky_retry_daemon(tmp_path_factory, self.crashes_before_success) + try: + yield daemon_info + finally: + stop_flaky_retry_daemon(daemon_info) diff --git a/feature_integration_tests/test_cases/support_apps/flaky_startup_app/BUILD b/feature_integration_tests/test_cases/support_apps/flaky_startup_app/BUILD new file mode 100644 index 00000000000..4eb147a7b69 --- /dev/null +++ b/feature_integration_tests/test_cases/support_apps/flaky_startup_app/BUILD @@ -0,0 +1,25 @@ +# ******************************************************************************* +# Copyright (c) 2026 Contributors to the Eclipse Foundation +# +# See the NOTICE file(s) distributed with this work for additional +# information regarding copyright ownership. +# +# This program and the accompanying materials are made available under the +# terms of the Apache License Version 2.0 which is available at +# https://www.apache.org/licenses/LICENSE-2.0 +# +# SPDX-License-Identifier: Apache-2.0 +# ******************************************************************************* +load("@rules_cc//cc:defs.bzl", "cc_binary") + +# "Reporting" supervised app (calls report_running() via lifecycle_cc) used only to +# deterministically drive ready_recovery_action.restart in +# lifecycle_daemon_retry_config.json. See main.cpp for behavior. +cc_binary( + name = "flaky_startup_app", + srcs = ["main.cpp"], + visibility = ["//feature_integration_tests:__subpackages__"], + deps = [ + "@score_lifecycle//score/launch_manager:lifecycle_cc", + ], +) diff --git a/feature_integration_tests/test_cases/support_apps/flaky_startup_app/main.cpp b/feature_integration_tests/test_cases/support_apps/flaky_startup_app/main.cpp new file mode 100644 index 00000000000..f20e6b0a6e0 --- /dev/null +++ b/feature_integration_tests/test_cases/support_apps/flaky_startup_app/main.cpp @@ -0,0 +1,100 @@ +/******************************************************************************** + * Copyright (c) 2026 Contributors to the Eclipse Foundation + * + * See the NOTICE file(s) distributed with this work for additional + * information regarding copyright ownership. + * + * This program and the accompanying materials are made available under the + * terms of the Apache License Version 2.0 which is available at + * https://www.apache.org/licenses/LICENSE-2.0 + * + * SPDX-License-Identifier: Apache-2.0 + ********************************************************************************/ + +// A "Reporting" launch_manager component that deterministically aborts on its +// first N startup attempts, then reports Running from attempt N+1 onward. The +// attempt count is persisted in a counter file so it survives across the +// process restarts that launch_manager performs in place, letting FITs +// exercise `ready_recovery_action.restart.number_of_attempts` (retry, and +// retry exhaustion) without relying on a real, racy startup failure. +// +// Must call report_running(): launch_manager's restart-in-place accounting is +// driven by the control-client channel that report_running() sets up. A plain +// "Native" process (no lifecycle API integration) never establishes that +// channel, so its startup failures go straight to the process group's +// recovery_action instead of being retried per `number_of_attempts`. + +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include + +namespace { + +std::atomic exit_requested{false}; + +void signal_handler(int signal) +{ + if (signal == SIGINT || signal == SIGTERM) + { + exit_requested = true; + } +} + +int read_attempt_count(const std::filesystem::path& counter_path) +{ + std::ifstream in(counter_path); + int count = 0; + if (in) + { + in >> count; + } + return count; +} + +void write_attempt_count(const std::filesystem::path& counter_path, int count) +{ + std::ofstream out(counter_path, std::ios::trunc); + out << count; +} + +} // namespace + +int main(int argc, char** argv) +{ + if (argc < 3) + { + std::cerr << "usage: flaky_startup_app " << std::endl; + return EXIT_FAILURE; + } + + const std::filesystem::path counter_path{argv[1]}; + const int crashes_before_success = std::atoi(argv[2]); + + const int attempt = read_attempt_count(counter_path); + write_attempt_count(counter_path, attempt + 1); + + if (attempt < crashes_before_success) + { + std::cerr << "flaky_startup_app: simulating crash on attempt " << attempt << std::endl; + std::abort(); + } + + std::cerr << "flaky_startup_app: starting successfully on attempt " << attempt << std::endl; + score::mw::lifecycle::report_running(); + + signal(SIGINT, signal_handler); + signal(SIGTERM, signal_handler); + while (!exit_requested) + { + std::this_thread::sleep_for(std::chrono::milliseconds(50)); + } + return EXIT_SUCCESS; +} diff --git a/feature_integration_tests/test_cases/tests/lifecycle/conftest.py b/feature_integration_tests/test_cases/tests/lifecycle/conftest.py new file mode 100644 index 00000000000..1d49444b948 --- /dev/null +++ b/feature_integration_tests/test_cases/tests/lifecycle/conftest.py @@ -0,0 +1,32 @@ +# ******************************************************************************* +# Copyright (c) 2026 Contributors to the Eclipse Foundation +# +# See the NOTICE file(s) distributed with this work for additional +# information regarding copyright ownership. +# +# This program and the accompanying materials are made available under the +# terms of the Apache License Version 2.0 which is available at +# https://www.apache.org/licenses/LICENSE-2.0 +# +# SPDX-License-Identifier: Apache-2.0 +# ******************************************************************************* +"""Shared fixtures for lifecycle daemon tests.""" + +from __future__ import annotations + +from typing import Any + +import pytest +from daemon_helpers import start_launch_manager_daemon, stop_launch_manager_daemon + + +@pytest.fixture(scope="class") +def launch_manager_daemon(tmp_path_factory: pytest.TempPathFactory) -> dict[str, Any]: + """One launch_manager (with both supervised apps running) per test class and `version` + param; see `daemon_helpers.start_launch_manager_daemon`. Tests using it must not start + another daemon.""" + daemon_info = start_launch_manager_daemon(tmp_path_factory) + try: + yield daemon_info + finally: + stop_launch_manager_daemon(daemon_info) diff --git a/feature_integration_tests/test_cases/tests/lifecycle/test_conditional_launching.py b/feature_integration_tests/test_cases/tests/lifecycle/test_conditional_launching.py new file mode 100644 index 00000000000..dcca057842c --- /dev/null +++ b/feature_integration_tests/test_cases/tests/lifecycle/test_conditional_launching.py @@ -0,0 +1,127 @@ +# ******************************************************************************* +# Copyright (c) 2026 Contributors to the Eclipse Foundation +# +# See the NOTICE file(s) distributed with this work for additional +# information regarding copyright ownership. +# +# This program and the accompanying materials are made available under the +# terms of the Apache License Version 2.0 which is available at +# https://www.apache.org/licenses/LICENSE-2.0 +# +# SPDX-License-Identifier: Apache-2.0 +# ******************************************************************************* +""" +Dependency-based launching FITs against a real launch_manager on lifecycle_daemon_config.json, +where rust_supervised_app `depends_on` cpp_supervised_app. (test_conditional_launching_scenario.py +only tests the FIT's own scenario stub.) +""" + +import json +from pathlib import Path +from typing import Any + +import pytest +from daemon_helpers import ( + is_running, + start_launch_manager_daemon, + stop_launch_manager_daemon, + wait_until, +) +from test_properties import add_test_properties + + +@pytest.mark.parametrize("version", ["rust", "cpp"], scope="class") +class TestConditionalLaunchingWithDaemon: + """Startup launch of each supervised app, per `version`, via the class-scoped fixture.""" + + @add_test_properties( + partially_verifies=["feat_req__lifecycle__launch_support"], + test_type="requirements-based", + derivation_technique="requirements-analysis", + ) + def test_startup_launches_conditioned_processes(self, launch_manager_daemon: dict[str, Any], version: str) -> None: + """The `version` app is running (pgrep) within 8 s of daemon start.""" + daemon_info = launch_manager_daemon + app_name = "rust_supervised_app" if version == "rust" else "cpp_supervised_app" + app_path = str(daemon_info["apps"][version]) + + started = wait_until(lambda: is_running(app_path), timeout_s=8.0) + assert started, f"{app_name} was not launched in conditional startup" + + +class TestConditionalLaunchingBlocksOnMissingDependency: + """rust startup is gated on cpp, observed by withholding cpp. + + Own daemon, in its own class so the class-scoped fixture is torn down first (fixed shm + names; see daemon_helpers._live_daemons). Not parametrized: runs once. + """ + + @add_test_properties( + partially_verifies=[ + "feat_req__lifecycle__waitfor_support", + "feat_req__lifecycle__dependency_check", + "feat_req__lifecycle__process_ordering", + "feat_req__lifecycle__define_swc_dependencies", + ], + test_type="requirements-based", + derivation_technique="requirements-analysis", + ) + def test_rust_stays_down_until_cpp_dependency_becomes_available( + self, tmp_path_factory: pytest.TempPathFactory + ) -> None: + """While cpp is non-executable, rust must stay down for 4 s and the daemon must log the cpp + launch failure at least twice (it keeps retrying rather than aborting). After cpp is made + executable, cpp and then rust must be running within 8 s each. + + Limitations: "running" is pgrep process existence; rust has a single dependency, so + `dependency_check` ("all dependencies") cannot be told apart from "any dependency". + No `cond_process_start` claim: that requirement is about starting on the return value + of earlier processes, which this config does not use. + """ + # Precondition: without this edge the negative check below would pass vacuously. + config_path = Path(__file__).resolve().parents[3] / "configs" / "lifecycle_daemon_config.json" + config = json.loads(config_path.read_text(encoding="utf-8")) + rust_depends_on = config["components"]["rust_supervised_app"]["component_properties"].get("depends_on", []) + assert "cpp_supervised_app" in rust_depends_on, ( + "Expected rust_supervised_app to depend on cpp_supervised_app in lifecycle daemon config" + ) + + daemon_info = start_launch_manager_daemon( + tmp_path_factory, + blocked_apps=frozenset({"cpp"}), + wait_for_apps=False, + ) + try: + cpp_path = daemon_info["apps"]["cpp"] + rust_path = str(daemon_info["apps"]["rust"]) + + # cpp is mode 0000, so its exec fails: rust must not appear meanwhile. + rust_started_early = wait_until(lambda: is_running(rust_path), timeout_s=4.0) + assert not rust_started_early, ( + "rust_supervised_app started even though its cpp_supervised_app dependency " + "was withheld (non-executable); dependency gating was not enforced" + ) + + # >= 2 path-specific launch failures: the daemon keeps retrying cpp, not giving up once. + cpp_launch_failure = f"File does not exist or is not executable: {cpp_path}" + failures_observed = wait_until( + lambda: daemon_info["daemon"].get_logs().count(cpp_launch_failure) >= 2, + 4.0, + ) + assert failures_observed, ( + "Expected repeated cpp launch failures while it was withheld; matching daemon logs:\n" + + "\n".join( + line for line in daemon_info["daemon"].get_logs().splitlines() if cpp_launch_failure in line + ) + ) + + # Unblock cpp: it, then its dependent rust, must start. + cpp_path.chmod(0o755) + cpp_started = wait_until(lambda: is_running(cpp_path), timeout_s=8.0) + assert cpp_started, "cpp_supervised_app did not start after becoming executable" + rust_started = wait_until(lambda: is_running(rust_path), timeout_s=8.0) + assert rust_started, ( + "rust_supervised_app did not start after its cpp_supervised_app dependency became available" + ) + finally: + stop_launch_manager_daemon(daemon_info) diff --git a/feature_integration_tests/test_cases/tests/lifecycle/test_conditional_launching_scenario.py b/feature_integration_tests/test_cases/tests/lifecycle/test_conditional_launching_scenario.py new file mode 100644 index 00000000000..b54d66eeeea --- /dev/null +++ b/feature_integration_tests/test_cases/tests/lifecycle/test_conditional_launching_scenario.py @@ -0,0 +1,356 @@ +# ******************************************************************************* +# Copyright (c) 2026 Contributors to the Eclipse Foundation +# +# See the NOTICE file(s) distributed with this work for additional +# information regarding copyright ownership. +# +# This program and the accompanying materials are made available under the +# terms of the Apache License Version 2.0 which is available at +# https://www.apache.org/licenses/LICENSE-2.0 +# +# SPDX-License-Identifier: Apache-2.0 +# ******************************************************************************* +"""Tests of the FIT's own conditional-launching scenario binary (rust and cpp), not of launch_manager. + +The scenario polls `path:`, `env:` and `process:` wait conditions; these tests really create or +withhold each condition and check what the stub reports. No test here drives launch_manager, so +none carries a `partially_verifies` claim: lifecycle requirement coverage lives in +test_conditional_launching.py and test_process_launching_with_daemon.py. +""" + +import os +import subprocess +import sys +import time +from collections.abc import Generator +from pathlib import Path +from typing import Any + +import pytest +from fit_scenario import ResultCode +from lifecycle_scenario import LifecycleScenario +from testing_utils import ScenarioResult + +pytestmark = [pytest.mark.parametrize("version", ["rust", "cpp"], scope="class")] + +_CONDITION_ENV_VAR = "LM_CONDITION_READY" +_CONDITION_PROCESS_NAME = "sleep" + + +class TestConditionalLaunchingScenario(LifecycleScenario): + """All three conditions are really satisfied before the scenario starts.""" + + @pytest.fixture(scope="class") + def scenario_name(self) -> str: + return "lifecycle.conditional_launching" + + @pytest.fixture(scope="class") + def flag_path(self, temp_dir: Path) -> Path: + return temp_dir / "lifecycle_launch_ready.flag" + + @pytest.fixture(scope="class", autouse=True) + def satisfied_preconditions(self, flag_path: Path) -> Generator[None, None, None]: + """Create the flag file, set the env var (inherited by the scenario) and keep a `sleep` + process alive for the class; all three are undone afterwards.""" + flag_path.write_text("ready", encoding="utf-8") + os.environ[_CONDITION_ENV_VAR] = "1" + process = subprocess.Popen([_CONDITION_PROCESS_NAME, "30"]) + try: + yield + finally: + process.kill() + process.wait() + del os.environ[_CONDITION_ENV_VAR] + flag_path.unlink(missing_ok=True) + + @pytest.fixture(scope="class") + def test_config(self, flag_path: Path, satisfied_preconditions: None) -> dict[str, Any]: + # Explicit dependency so the preconditions exist before `results` runs the scenario. + return { + "test": { + "wait_conditions": [ + f"path:{flag_path}", + f"env:{_CONDITION_ENV_VAR}", + f"process:{_CONDITION_PROCESS_NAME}", + ], + "polling_interval_ms": 50, + "timeout_ms": 2000, + }, + } + + def test_conditions_already_satisfied_allow_immediate_success( + self, + results: ScenarioResult, + version: str, + ) -> None: + """The scenario exits successfully.""" + assert results.return_code == ResultCode.SUCCESS, ( + f"Expected success with satisfied preconditions, got: {results}" + ) + + def test_each_condition_is_individually_confirmed_satisfied( + self, + logs_info_level: Any, + flag_path: Path, + version: str, + ) -> None: + """The scenario logs "Condition satisfied" for each condition and "All dependencies satisfied".""" + expected_messages = [ + f"Condition satisfied: path:{flag_path}", + f"Condition satisfied: env:{_CONDITION_ENV_VAR}", + f"Condition satisfied: process:{_CONDITION_PROCESS_NAME}", + "All dependencies satisfied", + ] + for expected in expected_messages: + log = logs_info_level.find_log("message", value=expected) + assert log is not None, f"Expected scenario to log: {expected}" + + def test_timeout_and_polling_interval_are_logged( + self, + logs_info_level: Any, + version: str, + ) -> None: + """The scenario logs the configured polling interval and timeout. Only the logged values + are checked; that polling actually uses them is covered by the late-condition test.""" + assert logs_info_level.find_log("message", value="Polling interval: 50ms") is not None + assert logs_info_level.find_log("message", value="Condition timeout: 2000ms") is not None + + +class TestConditionalLaunchingScenarioTimesOutOnUnmetConditions(LifecycleScenario): + """No condition is ever satisfied: the scenario must fail with a timeout. Catches a stub that + reports success without checking the conditions.""" + + @pytest.fixture(scope="class") + def scenario_name(self) -> str: + return "lifecycle.conditional_launching" + + @pytest.fixture(scope="class") + def test_config(self, temp_dir: Path) -> dict[str, Any]: + missing_path = temp_dir / "never_created.flag" + return { + "test": { + "wait_conditions": [ + f"path:{missing_path}", + "env:LM_CONDITION_NEVER_SET", + "process:process_that_does_not_exist_anywhere", + ], + "polling_interval_ms": 20, + "timeout_ms": 200, + }, + } + + def expect_command_failure(self) -> bool: + return True + + def capture_stderr(self) -> bool: + return True + + def test_scenario_fails_when_conditions_stay_unmet(self, results: ScenarioResult, version: str) -> None: + """Non-success exit with a wait-condition timeout ("Timed out" ... "condition") on stderr.""" + assert results.return_code != ResultCode.SUCCESS, ( + f"Expected failure when wait conditions are never satisfied, got: {results}" + ) + assert results.stderr is not None + assert "Timed out" in results.stderr and "condition" in results.stderr, ( + f"Expected a wait-condition timeout error on stderr, got: {results.stderr}" + ) + + +class TestConditionalLaunchingScenarioDetectsConditionArrivingLate(LifecycleScenario): + """The path condition becomes true 0.5 s into a 3 s wait. Catches a stub that checks only once + (at start or at timeout) instead of polling.""" + + _DELAY_BEFORE_CONDITION_MET_S = 0.5 + _TIMEOUT_MS = 3000 + _POLLING_INTERVAL_MS = 100 + + @pytest.fixture(scope="class") + def scenario_name(self) -> str: + return "lifecycle.conditional_launching" + + @pytest.fixture(scope="class") + def flag_path(self, temp_dir: Path) -> Path: + return temp_dir / "lifecycle_launch_ready_late.flag" + + @pytest.fixture(scope="class") + def test_config(self, flag_path: Path) -> dict[str, Any]: + return { + "test": { + "wait_conditions": [f"path:{flag_path}"], + "polling_interval_ms": self._POLLING_INTERVAL_MS, + "timeout_ms": self._TIMEOUT_MS, + }, + } + + @pytest.fixture(scope="class") + def results( + self, + command: list[str], + execution_timeout: float, + flag_path: Path, + ) -> Generator[ScenarioResult, None, None]: + # Overrides the base `results` fixture so the delayed flag writer starts together with the + # scenario; the autouse report fixture runs `results` before the test body, so arming it + # in the test would be too late. A helper subprocess (not a Python timer thread) writes + # the flag, so the delay does not depend on the test runner's thread scheduling. + start = time.monotonic() + trigger = subprocess.Popen( + [ + sys.executable, + "-c", + ( + "import pathlib, sys, time; " + "time.sleep(float(sys.argv[1])); " + "pathlib.Path(sys.argv[2]).write_text('ready', encoding='utf-8')" + ), + str(self._DELAY_BEFORE_CONDITION_MET_S), + str(flag_path), + ] + ) + try: + result = self._run_command(command, execution_timeout) + # Fixture and test get different instances, so share the timing via the class. + type(self)._elapsed_s = time.monotonic() - start + finally: + trigger.terminate() + try: + trigger.wait(timeout=1) + except subprocess.TimeoutExpired: + trigger.kill() + trigger.wait(timeout=1) + yield result + flag_path.unlink(missing_ok=True) + + def test_condition_satisfied_partway_through_the_wait_is_detected_promptly( + self, + results: ScenarioResult, + version: str, + ) -> None: + """Success no earlier than the 0.5 s delay and well before the 3 s timeout.""" + result = results + elapsed_s = self._elapsed_s + + assert result.return_code == ResultCode.SUCCESS, f"Expected success once the condition became true: {result}" + assert elapsed_s >= self._DELAY_BEFORE_CONDITION_MET_S, ( + f"Scenario reported success ({elapsed_s:.2f}s) before the condition could possibly have been " + f"true ({self._DELAY_BEFORE_CONDITION_MET_S}s) - it isn't actually observing the real condition." + ) + margin_s = self._TIMEOUT_MS / 1000 - self._DELAY_BEFORE_CONDITION_MET_S + assert elapsed_s < self._DELAY_BEFORE_CONDITION_MET_S + margin_s / 2, ( + f"Scenario took {elapsed_s:.2f}s to detect a condition that became true after " + f"{self._DELAY_BEFORE_CONDITION_MET_S}s - this is close to the full {self._TIMEOUT_MS}ms timeout, " + "suggesting it isn't re-checking at the configured polling interval." + ) + + +class TestConditionalLaunchingScenarioRejectsUnsupportedPrefix(LifecycleScenario): + """An unknown wait-condition prefix is a configuration error, reported without waiting.""" + + @pytest.fixture(scope="class") + def scenario_name(self) -> str: + return "lifecycle.conditional_launching" + + _TIMEOUT_MS = 2000 + + @pytest.fixture(scope="class") + def test_config(self) -> dict[str, Any]: + return { + "test": { + "wait_conditions": ["badprefix:value"], + "polling_interval_ms": 20, + "timeout_ms": self._TIMEOUT_MS, + }, + } + + def expect_command_failure(self) -> bool: + return True + + def capture_stderr(self) -> bool: + return True + + @pytest.fixture(scope="class") + def results(self, command: list[str], execution_timeout: float) -> ScenarioResult: + # Overrides the base `results` fixture to time the one scenario run (a second run in the + # test body would execute the binary twice). Fixture and test get different instances, + # so the timing is shared via the class. + start = time.monotonic() + result = self._run_command(command, execution_timeout) + type(self)._elapsed_s = time.monotonic() - start + return result + + def test_unsupported_prefix_is_rejected_immediately( + self, + results: ScenarioResult, + version: str, + ) -> None: + """Non-success exit with "Unsupported wait condition prefix" on stderr, in under half the timeout.""" + result = results + elapsed_s = self._elapsed_s + + assert result.return_code != ResultCode.SUCCESS, ( + f"Expected failure for an unsupported wait-condition prefix, got: {result}" + ) + assert result.stderr is not None + assert "Unsupported wait condition prefix" in result.stderr, ( + f"Expected an unsupported-prefix validation error on stderr, got: {result.stderr}" + ) + assert elapsed_s < (self._TIMEOUT_MS / 1000) / 2, ( + f"Rejection took {elapsed_s:.2f}s, close to the full {self._TIMEOUT_MS}ms timeout - " + "the prefix should be validated up front, not discovered by waiting it out." + ) + + +class TestConditionalLaunchingScenarioRejectsEmptyConditions(LifecycleScenario): + """An empty `wait_conditions` list is a configuration error (not "nothing to wait for"), + reported without waiting.""" + + @pytest.fixture(scope="class") + def scenario_name(self) -> str: + return "lifecycle.conditional_launching" + + _TIMEOUT_MS = 2000 + + @pytest.fixture(scope="class") + def test_config(self) -> dict[str, Any]: + return { + "test": { + "wait_conditions": [], + "polling_interval_ms": 20, + "timeout_ms": self._TIMEOUT_MS, + }, + } + + def expect_command_failure(self) -> bool: + return True + + def capture_stderr(self) -> bool: + return True + + @pytest.fixture(scope="class") + def results(self, command: list[str], execution_timeout: float) -> ScenarioResult: + # Same override as TestConditionalLaunchingScenarioRejectsUnsupportedPrefix.results. + start = time.monotonic() + result = self._run_command(command, execution_timeout) + type(self)._elapsed_s = time.monotonic() - start + return result + + def test_empty_conditions_are_rejected_immediately( + self, + results: ScenarioResult, + version: str, + ) -> None: + """Non-success exit with "Wait conditions were not provided" on stderr, in under half the timeout.""" + result = results + elapsed_s = self._elapsed_s + + assert result.return_code != ResultCode.SUCCESS, ( + f"Expected failure for an empty wait_conditions list, got: {result}" + ) + assert result.stderr is not None + assert "Wait conditions were not provided" in result.stderr, ( + f"Expected a missing/empty wait_conditions validation error on stderr, got: {result.stderr}" + ) + assert elapsed_s < (self._TIMEOUT_MS / 1000) / 2, ( + f"Rejection took {elapsed_s:.2f}s, close to the full {self._TIMEOUT_MS}ms timeout - " + "an empty condition list should be validated up front, not discovered by waiting it out." + ) diff --git a/feature_integration_tests/test_cases/tests/lifecycle/test_process_launching_with_daemon.py b/feature_integration_tests/test_cases/tests/lifecycle/test_process_launching_with_daemon.py new file mode 100644 index 00000000000..1862d530f51 --- /dev/null +++ b/feature_integration_tests/test_cases/tests/lifecycle/test_process_launching_with_daemon.py @@ -0,0 +1,569 @@ +# ******************************************************************************* +# Copyright (c) 2026 Contributors to the Eclipse Foundation +# +# See the NOTICE file(s) distributed with this work for additional +# information regarding copyright ownership. +# +# This program and the accompanying materials are made available under the +# terms of the Apache License Version 2.0 which is available at +# https://www.apache.org/licenses/LICENSE-2.0 +# +# SPDX-License-Identifier: Apache-2.0 +# ******************************************************************************* +""" +Lifecycle FITs against a real launch_manager supervising rust_supervised_app and cpp_supervised_app. + +`version` selects which of the two supervised apps a test inspects; both always run under the +same daemon. Run via Bazel (the helpers resolve binaries from the target's FIT_*_PATH env vars): + + bazel test //feature_integration_tests/test_cases:fit_lifecycle_daemon + bazel test //feature_integration_tests/test_cases:fit_lifecycle_daemon --test_arg=-k --test_arg=rust +""" + +import json +import os +import re +import subprocess +import time +from pathlib import Path +from typing import Any + +import pytest +from daemon_helpers import ( + first_pid, + is_running, + read_proc_file, + signal_process, + start_launch_manager_daemon, + stop_launch_manager_daemon, + wait_until, +) +from test_properties import add_test_properties + +pytestmark = [ + pytest.mark.parametrize("version", ["rust", "cpp"], scope="class"), +] + + +class TestProcessLaunchingWithDaemon: + """Launch-parameter checks (args, env, uid/gid, scheduling, non-root) against one daemon per + `version`, provided by the class-scoped `launch_manager_daemon` fixture. Tests here must not + start their own daemon (fixed shm names; see daemon_helpers._live_daemons).""" + + @staticmethod + def _proc_cmdline(pid: str) -> list[str]: + """Return `/proc//cmdline` as a list of arguments.""" + raw = Path(f"/proc/{pid}/cmdline").read_bytes() + return [arg.decode("utf-8") for arg in raw.split(b"\0") if arg] + + @staticmethod + def _proc_environ(pid: str) -> dict[str, str]: + """Return `/proc//environ` as a dict (via the privileged `cat` when staged).""" + raw = read_proc_file(pid, "environ") + env: dict[str, str] = {} + for item in raw.split(b"\0"): + if not item: + continue + key, sep, value = item.partition(b"=") + if not sep: + continue + env[key.decode("utf-8")] = value.decode("utf-8") + return env + + @staticmethod + def _proc_status_ids(pid: str) -> tuple[int, int] | None: + """Return the effective `(uid, gid)` from `/proc//status`, or None if unreadable.""" + try: + lines = Path(f"/proc/{pid}/status").read_text(encoding="utf-8").splitlines() + except OSError: + return None + uid_line = next((line for line in lines if line.startswith("Uid:")), None) + gid_line = next((line for line in lines if line.startswith("Gid:")), None) + if uid_line is None or gid_line is None: + return None + try: + uid_parts = uid_line.split()[1:] + gid_parts = gid_line.split()[1:] + # /proc status format: real effective saved filesystem + return int(uid_parts[1]), int(gid_parts[1]) + except (IndexError, ValueError): + return None + + @staticmethod + def _proc_sched_policy_and_priority(pid: str) -> tuple[str, int] | None: + """Return `(policy, priority)` parsed from `chrt -p `, or None on failure.""" + result = subprocess.run(["chrt", "-p", pid], capture_output=True, text=True, check=False) + if result.returncode != 0: + return None + + policy = None + priority = None + for line in result.stdout.splitlines(): + lower = line.lower().strip() + # e.g. "pid 400132's current scheduling policy: SCHED_OTHER" + if "scheduling policy" in lower: + policy = line.split(":", 1)[1].strip() + elif "scheduling priority" in lower: + try: + priority = int(line.split(":", 1)[1].strip()) + except ValueError: + return None + + if policy is None or priority is None: + return None + return policy, priority + + # Dependency gating (rust waits for cpp) is covered in test_conditional_launching.py. + + @add_test_properties( + partially_verifies=["feat_req__lifecycle__launch_support"], + test_type="requirements-based", + derivation_technique="requirements-analysis", + ) + def test_startup_declares_and_launches_multiple_processes( + self, + launch_manager_daemon: dict[str, Any], + version: str, + ) -> None: + """Both processes in the Startup run target's `depends_on` are running (pgrep) and the + daemon is still up. `version` is unused: the check covers both apps.""" + config_path = Path(__file__).resolve().parents[3] / "configs" / "lifecycle_daemon_config.json" + config = json.loads(config_path.read_text(encoding="utf-8")) + startup_deps = config["run_targets"]["Startup"]["depends_on"] + + assert isinstance(startup_deps, list), "Startup depends_on should be a list" + assert len(startup_deps) >= 2, "Startup run target should define multiple process dependencies" + assert "cpp_supervised_app" in startup_deps, "cpp_supervised_app missing in Startup depends_on" + assert "rust_supervised_app" in startup_deps, "rust_supervised_app missing in Startup depends_on" + + daemon_info = launch_manager_daemon + daemon = daemon_info["daemon"] + cpp_path = str(daemon_info["apps"]["cpp"]) + rust_path = str(daemon_info["apps"]["rust"]) + both_running = wait_until( + lambda: is_running(cpp_path) and is_running(rust_path), + timeout_s=8.0, + ) + assert both_running, "Startup should launch all configured supervised processes" + assert daemon.is_running(), "Launch Manager daemon stopped unexpectedly" + + @add_test_properties( + partially_verifies=["feat_req__lifecycle__process_launch_args"], + test_type="requirements-based", + derivation_technique="requirements-analysis", + ) + def test_launch_process_arguments_are_applied( + self, + launch_manager_daemon: dict[str, Any], + version: str, + ) -> None: + """Every configured `process_arguments` entry is present in the live app's cmdline. + Expected values come from the config, so the test tracks config changes.""" + daemon_info = launch_manager_daemon + app_name = "rust_supervised_app" if version == "rust" else "cpp_supervised_app" + app_path = str(daemon_info["apps"][version]) + + config_path = Path(__file__).resolve().parents[3] / "configs" / "lifecycle_daemon_config.json" + config = json.loads(config_path.read_text(encoding="utf-8")) + configured_args = config["components"][app_name]["component_properties"]["process_arguments"] + assert configured_args, f"{app_name} does not configure any process_arguments to verify against" + + started = wait_until(lambda: is_running(app_path), timeout_s=8.0) + assert started, f"{app_name} was not launched before argument verification" + + pid = first_pid(app_path) + assert pid is not None, f"Could not resolve PID for {app_name}" + cmdline = self._proc_cmdline(pid) + + assert cmdline, f"Could not read command line arguments for {app_name} pid={pid}" + for configured_arg in configured_args: + assert configured_arg in cmdline, ( + f"Configured launch argument {configured_arg!r} missing in {app_name} cmdline: {cmdline}" + ) + + @add_test_properties( + partially_verifies=["feat_req__lifecycle__process_launch_args"], + test_type="requirements-based", + derivation_technique="requirements-analysis", + ) + def test_launch_process_environment_is_applied( + self, + launch_manager_daemon: dict[str, Any], + version: str, + ) -> None: + """Every configured `environmental_variables` entry has the configured value in the live + app's /proc environ. Expected values come from the config.""" + daemon_info = launch_manager_daemon + app_name = "rust_supervised_app" if version == "rust" else "cpp_supervised_app" + app_path = str(daemon_info["apps"][version]) + + config_path = Path(__file__).resolve().parents[3] / "configs" / "lifecycle_daemon_config.json" + config = json.loads(config_path.read_text(encoding="utf-8")) + configured_env = config["components"][app_name]["deployment_config"]["environmental_variables"] + assert configured_env, f"{app_name} does not configure any environmental_variables to verify against" + + started = wait_until(lambda: is_running(app_path), timeout_s=8.0) + assert started, f"{app_name} was not launched before environment verification" + + pid = first_pid(app_path) + assert pid is not None, f"Could not resolve PID for {app_name}" + proc_env = self._proc_environ(pid) + + for key, expected_value in configured_env.items(): + assert proc_env.get(key) == expected_value, ( + f"{key} mismatch for {app_name}: expected {expected_value!r}, got {proc_env.get(key)!r}" + ) + + # No requirement claim (feat_req__lifecycle__uid_gid_support): skips in CI, which neither sets + # FIT_ENABLE_SETCAP=1 nor runs unsandboxed, both needed for the setcap grant. See README.md. + def test_launched_process_uid_gid_matches_config_when_applied( + self, + launch_manager_daemon: dict[str, Any], + version: str, + ) -> None: + """With capabilities granted, the live app's effective uid equals the rendered sandbox uid + and differs from the runner's. Skips without the grant, or if the uid was not remapped. + + Limitation: the rendered gid is the runner's own (see `_generate_runtime_config`), so the + gid assertion cannot prove setgid() was applied. + """ + daemon_info = launch_manager_daemon + if not daemon_info["sandbox_privileged"]: + pytest.skip( + "launch_manager was not granted cap_setuid/cap_setgid in this environment; " + f"sandbox uid/gid cannot be applied. Reason: {daemon_info['sandbox_privileged_reason']}" + ) + + app_name = "rust_supervised_app" if version == "rust" else "cpp_supervised_app" + app_path = str(daemon_info["apps"][version]) + + # The rendered config, not the source JSON: under capabilities the helper remaps the + # sandbox uid away from the runner's own (daemon_helpers._SANDBOX_UID). + config = json.loads(daemon_info["runtime_config"].read_text(encoding="utf-8")) + component_sandbox = config["components"][app_name].get("deployment_config", {}).get("sandbox") + sandbox = component_sandbox or config["defaults"]["deployment_config"]["sandbox"] + expected_uid = int(sandbox["uid"]) + expected_gid = int(sandbox["gid"]) + if expected_uid == os.getuid(): + # Happens when the kill/cat copies could not be staged (uid not remapped); matching uids pass + # whether or not launch_manager applied the sandbox identity. + pytest.skip( + f"Sandbox uid {expected_uid} equals the test runner's own uid; cannot distinguish an " + f"applied sandbox identity from an inherited one. Reason: {daemon_info['sandbox_privileged_reason']}" + ) + + started = wait_until(lambda: is_running(app_path), timeout_s=8.0) + assert started, f"{app_name} was not launched before uid/gid verification" + + pid = first_pid(app_path) + assert pid is not None, f"Could not resolve PID for {app_name}" + proc_ids = self._proc_status_ids(pid) + assert proc_ids is not None, f"Could not read /proc status uid/gid for {app_name} pid={pid}" + + effective_uid, effective_gid = proc_ids + assert effective_uid == expected_uid, ( + f"Effective uid mismatch for {app_name}: expected {expected_uid}, got {effective_uid}" + ) + # Consistency check only: passes whether or not setgid() was applied (gid == runner's). + assert effective_gid == expected_gid, ( + f"Effective gid mismatch for {app_name}: expected {expected_gid}, got {effective_gid}" + ) + assert effective_uid != os.getuid(), ( + f"{app_name} is running as the test runner's own uid ({effective_uid}); sandbox " + "identity was not actually applied" + ) + + # No requirement claim (launch_priority_support / scheduling_policy): skips in CI, see above. + def test_launched_process_scheduling_matches_config_when_applied( + self, + launch_manager_daemon: dict[str, Any], + version: str, + ) -> None: + """With capabilities granted, the live app's `chrt` policy and priority equal the rendered + sandbox values. Skips without the grant (the config is then downgraded to SCHED_OTHER).""" + daemon_info = launch_manager_daemon + if not daemon_info["sandbox_privileged"]: + pytest.skip( + "launch_manager was not granted cap_sys_nice in this environment; " + f"scheduling policy cannot be applied. Reason: {daemon_info['sandbox_privileged_reason']}" + ) + + app_name = "rust_supervised_app" if version == "rust" else "cpp_supervised_app" + app_path = str(daemon_info["apps"][version]) + + # The rendered config, not the source JSON, like the uid/gid test above. + config = json.loads(daemon_info["runtime_config"].read_text(encoding="utf-8")) + component_sandbox = config["components"][app_name].get("deployment_config", {}).get("sandbox") + sandbox = component_sandbox or config["defaults"]["deployment_config"]["sandbox"] + configured_policy = sandbox["scheduling_policy"] + configured_priority = int(sandbox["scheduling_priority"]) + + started = wait_until(lambda: is_running(app_path), timeout_s=8.0) + assert started, f"{app_name} was not launched before scheduling verification" + + # Retry with a fresh pid in case the app restarted between pgrep and chrt. + sched = None + pid = None + for _ in range(20): + pid = first_pid(app_path) + if pid is None: + time.sleep(0.1) + continue + sched = self._proc_sched_policy_and_priority(pid) + if sched is not None: + break + time.sleep(0.1) + assert pid is not None, f"Could not resolve PID for {app_name}" + assert sched is not None, f"Could not read scheduling metadata via chrt for {app_name} pid={pid}" + + policy, rt_priority = sched + expected_policy = configured_policy.upper() + assert policy.upper() == expected_policy, ( + f"Scheduling policy mismatch for {app_name}: expected {expected_policy}, got {policy}" + ) + assert rt_priority == configured_priority, ( + f"Scheduling priority mismatch for {app_name}: expected {configured_priority}, got {rt_priority}" + ) + + # No requirement claim: skips in CI, see above. + def test_scheduling_policy_is_non_default_and_applied( + self, + launch_manager_daemon: dict[str, Any], + version: str, + ) -> None: + """With capabilities granted, the live app's `(policy, priority)` differs from + launch_manager's own, so it was set by launch_manager rather than inherited. The config + gives rust SCHED_RR/10 and cpp SCHED_FIFO/20; the daemon runs SCHED_OTHER/0. Skips + without the grant.""" + daemon_info = launch_manager_daemon + if not daemon_info["sandbox_privileged"]: + pytest.skip( + "launch_manager was not granted cap_sys_nice in this environment; " + f"scheduling policy cannot be applied. Reason: {daemon_info['sandbox_privileged_reason']}" + ) + + app_name = "rust_supervised_app" if version == "rust" else "cpp_supervised_app" + app_path = str(daemon_info["apps"][version]) + + started = wait_until(lambda: is_running(app_path), timeout_s=8.0) + assert started, f"{app_name} was not launched before scheduling verification" + + daemon_sched = self._proc_sched_policy_and_priority(str(daemon_info["daemon"].pid())) + assert daemon_sched is not None, "Could not read scheduling metadata via chrt for launch_manager" + + pid = None + app_sched = None + for _ in range(20): + pid = first_pid(app_path) + if pid is None: + time.sleep(0.1) + continue + app_sched = self._proc_sched_policy_and_priority(pid) + if app_sched is not None: + break + time.sleep(0.1) + assert pid is not None, f"Could not resolve PID for {app_name}" + assert app_sched is not None, f"Could not read scheduling metadata via chrt for {app_name} pid={pid}" + + assert app_sched != daemon_sched, ( + f"{app_name}'s scheduling {app_sched} matches launch_manager's own {daemon_sched}; " + "sandbox scheduling policy was not actually applied" + ) + + @add_test_properties( + partially_verifies=["feat_req__lifecycle__secpol_non_root"], + test_type="requirements-based", + derivation_technique="requirements-analysis", + ) + def test_launch_manager_and_apps_are_not_running_as_root( + self, + launch_manager_daemon: dict[str, Any], + version: str, + ) -> None: + """launch_manager itself runs with a non-root effective uid and still launches the app + (the requirement: LM can be started as non-root). Also checks the app is non-root. + + Limitation: covers only a plain non-root start (optionally with file capabilities), not + any other "security policy" mechanism. + """ + daemon_info = launch_manager_daemon + daemon = daemon_info["daemon"] + app_name = "rust_supervised_app" if version == "rust" else "cpp_supervised_app" + app_path = str(daemon_info["apps"][version]) + + assert daemon.is_running(), "Launch Manager daemon is not running" + daemon_ids = self._proc_status_ids(str(daemon.pid())) + assert daemon_ids is not None, f"Could not read /proc status for launch_manager pid={daemon.pid()}" + assert daemon_ids[0] != 0, "launch_manager is running as root" + + started = wait_until(lambda: is_running(app_path), timeout_s=8.0) + assert started, f"{app_name} was not launched before non-root verification" + + pid = first_pid(app_path) + assert pid is not None, f"Could not resolve PID for {app_name}" + proc_ids = self._proc_status_ids(pid) + assert proc_ids is not None, f"Could not read /proc status uid/gid for {app_name} pid={pid}" + effective_uid, _ = proc_ids + assert effective_uid != 0, f"{app_name} is unexpectedly running as root" + + +class TestSupervisedAppRecovery: + """Kill-and-recover against a dedicated daemon per `version`. + + Separate class so its daemon never overlaps the `launch_manager_daemon` fixture (fixed shm + names; see daemon_helpers._live_daemons). + """ + + @add_test_properties( + partially_verifies=[ + "feat_req__lifecycle__monitor_abnormal_term", + "feat_req__lifecycle__recovery_action_support", + ], + test_type="requirements-based", + derivation_technique="requirements-analysis", + ) + def test_supervised_app_recovery( + self, + tmp_path_factory: pytest.TempPathFactory, + version: str, + ) -> None: + """SIGKILL the running app; launch_manager must log its unexpected termination and bring + it back (new pid) without restarting the other, healthy app or dying itself. + + The recovery that runs is the run target's `recovery_action` (switch to + `fallback_run_target`, which contains both apps), not `ready_recovery_action.restart`, + which only covers startup failures. Limitation: the test does not assert which recovery + action ran, only that the app recovered. Does not claim `retries_configurable` + (see test_retry_exhaustion.py). + """ + daemon_info = start_launch_manager_daemon(tmp_path_factory) + try: + daemon = daemon_info["daemon"] + app_name = "rust_supervised_app" if version == "rust" else "cpp_supervised_app" + app_path = str(daemon_info["apps"][version]) + other_version = "cpp" if version == "rust" else "rust" + other_app_path = str(daemon_info["apps"][other_version]) + + started = wait_until(lambda: is_running(app_path), timeout_s=8.0) + assert started, f"{app_name} was not running before recovery test" + + old_pid = first_pid(app_path) + assert old_pid is not None, f"Could not resolve PID for {app_name}" + other_old_pid = first_pid(other_app_path) + assert other_old_pid is not None, "Could not resolve PID for the other supervised app" + + sent, reason = signal_process(old_pid, "-9", sandbox_privileged=daemon_info["sandbox_privileged"]) + assert sent, f"Could not signal {app_name} (pid={old_pid}): {reason}" + + restarted = wait_until( + lambda: (new_pid := first_pid(app_path)) is not None and new_pid != old_pid, + timeout_s=12.0, + ) + assert restarted, f"{app_name} was not restarted after forced termination" + termination_logged = re.search( + rf"unexpected termination of process\s+{re.escape(app_name)}\b", daemon.get_logs() + ) + assert termination_logged, ( + f"launch_manager did not log the abnormal termination of {app_name}.\nDaemon logs:\n{daemon.get_logs()}" + ) + + assert daemon.is_running(), "Launch Manager daemon should still be running after recovery" + + other_new_pid = first_pid(other_app_path) + assert other_new_pid == other_old_pid, ( + "The other, healthy supervised app was relaunched too; recovery should only relaunch the failed app" + ) + finally: + stop_launch_manager_daemon(daemon_info) + + +class TestParallelLaunch: + """Parallel launch of independent components, with its own daemons rendered with + `independent_apps=True`. `version` is unused (module-level parametrize), so the test runs + twice with identical behaviour.""" + + @add_test_properties( + partially_verifies=["feat_req__lifecycle__parallel_launch_support"], + test_type="requirements-based", + derivation_technique="requirements-analysis", + ) + def test_independent_processes_launch_without_waiting_on_each_other( + self, + tmp_path_factory: pytest.TempPathFactory, + version: str, + ) -> None: + """With no `depends_on` between the apps and one of them stalled (runs, never reports + Running), the other must be running within 4 s of the 1 s startup window. A serialized + launcher would wait out `ready_timeout` (10 s) on the stalled app first. Both stall + orders are tried, and the stalled stub must also be running (it was launched, not skipped). + """ + ready_timeout_s = 10.0 # rendered over the base config's 2.0 s + parallel_window_s = 4.0 # plus the 1 s startup window, still well under ready_timeout + assert parallel_window_s < ready_timeout_s / 2 + for stalled, other in (("cpp", "rust"), ("rust", "cpp")): + daemon_info = start_launch_manager_daemon( + tmp_path_factory, + stalled_apps=frozenset({stalled}), + wait_for_apps=False, + independent_apps=True, + ready_timeout_s=ready_timeout_s, + ) + try: + stalled_path = str(daemon_info["apps"][stalled]) + other_path = str(daemon_info["apps"][other]) + other_started = wait_until(lambda p=other_path: is_running(p), timeout_s=parallel_window_s) + assert other_started, ( + f"{other}_supervised_app did not start within {parallel_window_s}s while " + f"{stalled}_supervised_app was stalled (ready_timeout={ready_timeout_s}s), even " + "though neither depends on the other - launch is serialized, not parallel" + ) + # Both in flight at once: rules out the stalled stub never being launched. + assert is_running(stalled_path), f"stalled {stalled}_supervised_app stub was never launched" + finally: + stop_launch_manager_daemon(daemon_info) + + +class TestHealthMonitoringWithDaemon: + """Alive-supervision (watchdog) detection, with its own daemon per `version`.""" + + # No `smart_watchdog_config` claim: alive supervision is configured only in `defaults`, and + # the test neither configures it per process nor varies it. + @add_test_properties( + partially_verifies=["feat_req__lifecycle__liveliness_detection"], + test_type="requirements-based", + derivation_technique="requirements-analysis", + ) + def test_watchdog_detection(self, tmp_path_factory: pytest.TempPathFactory, version: str) -> None: + """SIGSTOP the app so it stops reporting alive indications; launch_manager must log its + Alive Supervision switching to FAILED or EXPIRED within 8 s. + + Limitation: checks detection only, not the reaction that follows. Uses its own daemon + so the recovery triggered by one version's failure cannot affect the other's run. + """ + daemon_info = start_launch_manager_daemon(tmp_path_factory) + try: + self._check_watchdog_detection(daemon_info, version) + finally: + stop_launch_manager_daemon(daemon_info) + + @staticmethod + def _check_watchdog_detection(daemon_info: dict[str, Any], version: str) -> None: + daemon = daemon_info["daemon"] + app_name = "rust_supervised_app" if version == "rust" else "cpp_supervised_app" + + app_path = str(daemon_info["apps"][version]) + # start_launch_manager_daemon already waited for the app, so a missing process is a failure. + pid = first_pid(app_path) + assert pid is not None, f"{app_name} died before the watchdog check" + + sandbox_privileged = daemon_info["sandbox_privileged"] + sent, reason = signal_process(pid, "-STOP", sandbox_privileged=sandbox_privileged) + assert sent, f"Could not signal {app_name} (pid={pid}): {reason}" + try: + # Only alive-supervision verdicts count; a startup timeout or crash is not liveliness detection. + liveliness_lost = rf"Alive Supervision \(\s*{re.escape(app_name)}\s*\) switched to (FAILED|EXPIRED)" + # Poll rather than sleep: detection latency varies under CI load. + detected = wait_until(lambda: re.search(liveliness_lost, daemon.get_logs()), timeout_s=8.0) + assert detected, f"No Alive Supervision failure logged for {app_name}.\nDaemon logs:\n{daemon.get_logs()}" + finally: + signal_process(pid, "-CONT", sandbox_privileged=sandbox_privileged) diff --git a/feature_integration_tests/test_cases/tests/lifecycle/test_retry_exhaustion.py b/feature_integration_tests/test_cases/tests/lifecycle/test_retry_exhaustion.py new file mode 100644 index 00000000000..7f3f233f434 --- /dev/null +++ b/feature_integration_tests/test_cases/tests/lifecycle/test_retry_exhaustion.py @@ -0,0 +1,126 @@ +# ******************************************************************************* +# Copyright (c) 2026 Contributors to the Eclipse Foundation +# +# See the NOTICE file(s) distributed with this work for additional +# information regarding copyright ownership. +# +# This program and the accompanying materials are made available under the +# terms of the Apache License Version 2.0 which is available at +# https://www.apache.org/licenses/LICENSE-2.0 +# +# SPDX-License-Identifier: Apache-2.0 +# ******************************************************************************* +"""Startup-retry FITs for `ready_recovery_action.restart.number_of_attempts`. + +`flaky_startup_app` aborts on its first `crashes_before_success` launches (rendered per class) +and counts every launch in a file, so the daemon's retry count is observed exactly: + +- `TestRetrySucceedsWithinConfiguredAttempts`: crashes use up all retries, then the app runs. +- `TestRetryExhaustionTriggersRecovery`: the app always crashes; the daemon must stop after + 1 + `number_of_attempts` launches and stay alive. + +Limitation for `retries_configurable`: only the single configured value (2) is exercised; the +count is not varied across runs. +""" + +from __future__ import annotations + +import json +from typing import Any + +import pytest +from daemon_helpers import ( + is_running, + read_retry_attempt_count, + wait_until, +) +from lifecycle_scenario import RetryDaemonScenario +from test_properties import add_test_properties + + +def _number_of_attempts(retry_daemon: dict[str, Any]) -> int: + """flaky_startup_app's `number_of_attempts`, read from the rendered config.""" + config = json.loads(retry_daemon["runtime_config"].read_text(encoding="utf-8")) + deployment = config["components"]["flaky_startup_app"]["deployment_config"] + return int(deployment["ready_recovery_action"]["restart"]["number_of_attempts"]) + + +class TestRetrySucceedsWithinConfiguredAttempts(RetryDaemonScenario): + """The component crashes exactly `number_of_attempts` times, then succeeds.""" + + crashes_before_success = 2 + + @add_test_properties( + partially_verifies=["feat_req__lifecycle__retries_configurable"], + test_type="requirements-based", + derivation_technique="requirements-analysis", + ) + def test_component_recovers_within_configured_attempts(self, retry_daemon: dict[str, Any]) -> None: + """Exactly `crashes_before_success + 1` launches happen, the last one stays running (pgrep), + and no further launch follows within 1.5 s.""" + app_path = retry_daemon["app_path"] + counter_path = retry_daemon["counter_path"] + expected_attempts = retry_daemon["crashes_before_success"] + 1 + assert retry_daemon["crashes_before_success"] <= _number_of_attempts(retry_daemon), ( + "crashes_before_success must fit within number_of_attempts for this test to be meaningful" + ) + + # Wait on the counter, not is_running(): a crashing attempt is briefly visible to pgrep. + reached = wait_until(lambda: read_retry_attempt_count(counter_path) >= expected_attempts, timeout_s=8.0) + assert reached, "flaky_startup_app never reached the expected number of launch attempts" + + attempts = read_retry_attempt_count(counter_path) + assert attempts == expected_attempts, ( + f"Expected exactly {expected_attempts} launch attempts (crashes_before_success + 1 success), got {attempts}" + ) + + started = wait_until(lambda: is_running(app_path), timeout_s=2.0) + assert started, "flaky_startup_app never reached Running after its last launch attempt" + + relaunched = wait_until(lambda: read_retry_attempt_count(counter_path) != attempts, timeout_s=1.5) + assert not relaunched, "Component was relaunched again after it was already Running" + assert is_running(app_path), "flaky_startup_app stopped running after recovering" + + +class TestRetryExhaustionTriggersRecovery(RetryDaemonScenario): + """The component always crashes, exhausting `number_of_attempts`.""" + + crashes_before_success = 999 + + @add_test_properties( + partially_verifies=["feat_req__lifecycle__retries_configurable"], + test_type="requirements-based", + derivation_technique="requirements-analysis", + ) + def test_daemon_gives_up_after_configured_attempts(self, retry_daemon: dict[str, Any]) -> None: + """Exactly `1 + number_of_attempts` launches happen, none follows within 2 s, the app is + not running and the daemon is still up. + + Limitation: the run target's `recovery_action` (switch to `fallback_run_target`) is only + inferred from the daemon staying up without relaunching; it is not asserted directly. + """ + app_path = retry_daemon["app_path"] + counter_path = retry_daemon["counter_path"] + number_of_attempts = _number_of_attempts(retry_daemon) + + settled = wait_until( + lambda: read_retry_attempt_count(counter_path) >= number_of_attempts + 1, + timeout_s=8.0, + ) + assert settled, "flaky_startup_app never reached the configured number of launch attempts" + + # Would catch a daemon that keeps retrying past the budget. + attempts_after_exhaustion = read_retry_attempt_count(counter_path) + kept_retrying = wait_until( + lambda: read_retry_attempt_count(counter_path) != attempts_after_exhaustion, + timeout_s=2.0, + ) + assert not kept_retrying, ( + f"Daemon kept restarting the component past the configured number_of_attempts={number_of_attempts}" + ) + assert attempts_after_exhaustion == number_of_attempts + 1, ( + f"Expected exactly {number_of_attempts + 1} launch attempts before giving up, got " + f"{attempts_after_exhaustion}" + ) + assert not is_running(app_path), "flaky_startup_app is still running after exhausting its retries" + assert retry_daemon["daemon"].is_running(), "launch_manager exited after the component exhausted its retries" diff --git a/feature_integration_tests/test_scenarios/cpp/src/internals/log_helpers.h b/feature_integration_tests/test_scenarios/cpp/src/internals/log_helpers.h new file mode 100644 index 00000000000..5df9f60b11d --- /dev/null +++ b/feature_integration_tests/test_scenarios/cpp/src/internals/log_helpers.h @@ -0,0 +1,135 @@ +/******************************************************************************** + * Copyright (c) 2026 Contributors to the Eclipse Foundation + * + * See the NOTICE file(s) distributed with this work for additional + * information regarding copyright ownership. + * + * This program and the accompanying materials are made available under the + * terms of the Apache License Version 2.0 which is available at + * https://www.apache.org/licenses/LICENSE-2.0 + * + * SPDX-License-Identifier: Apache-2.0 + ********************************************************************************/ + +#ifndef INTERNALS_LOG_HELPERS_H_ +#define INTERNALS_LOG_HELPERS_H_ + +#include +#include +#include +#include +#include +#include + +namespace log_helpers { + +/** + * @brief Return the current UNIX timestamp as a decimal string (seconds). + * + * Used to populate the "timestamp" field in structured JSON log lines so that + * the C++ output matches the Rust tracing JSON shape expected by the FIT log + * filters. + * + * @return String containing the number of seconds since the UNIX epoch. + */ +inline std::string unix_seconds_string() { + const auto now = std::chrono::system_clock::now(); + const auto secs = + std::chrono::duration_cast(now.time_since_epoch()).count(); + return std::to_string(secs); +} + +/** + * @brief Escape a string for embedding as a JSON string value. + * + * Escapes '"', '\\', and control characters. Without this, interpolating a raw + * value (e.g. a filesystem path containing '"' or '\\') straight into a JSON + * string produces malformed JSON that downstream JSON-based log parsing (e.g. + * Python's FIT LogContainer) cannot read back. + * + * @param value Raw string to escape. + * @return JSON-escaped string, without surrounding quotes. + */ +inline std::string json_escape(const std::string& value) { + std::string escaped; + escaped.reserve(value.size()); + for (const char c : value) { + switch (c) { + case '"': + escaped += "\\\""; + break; + case '\\': + escaped += "\\\\"; + break; + case '\n': + escaped += "\\n"; + break; + case '\r': + escaped += "\\r"; + break; + case '\t': + escaped += "\\t"; + break; + default: + if (static_cast(c) < 0x20) { + std::ostringstream oss; + oss << "\\u" << std::hex << std::setfill('0') << std::setw(4) + << static_cast(static_cast(c)); + escaped += oss.str(); + } else { + escaped += c; + } + } + } + return escaped; +} + +/** + * @brief Emit a structured JSON INFO log line to stdout. + * + * Matches the Rust tracing JSON format expected by the FIT LogContainer so + * that Python test assertions can use find_log() uniformly for both Rust and + * C++ scenarios. + * + * Example output: + * @code + * {"timestamp":"1234567890","level":"INFO","fields":{"key":"my_key","value":42.0}, + * "target":"cpp_test_scenarios::scenarios::persistency::my_module","threadId":"ThreadId(1)"} + * @endcode + * + * @param fields JSON fragment for the "fields" object, e.g. @c "\"key\":\"x\",\"value\":1.0" + * Caller is responsible for escaping any string values embedded here. + * @param target Module target string embedded in the log line. + */ +inline void log_info(const std::string& fields, const std::string& target) { + std::cout << "{\"timestamp\":\"" << unix_seconds_string() + << "\",\"level\":\"INFO\",\"fields\":{" << fields + << "},\"target\":\"" << json_escape(target) + << "\",\"threadId\":\"ThreadId(1)\"}\n"; +} + +/** + * @brief Format a double value to match Python's str(float) representation. + * + * For whole-number values (e.g. 42.0, 200.0) this appends ".0" so that the + * resulting string matches what Python's f-string interpolation produces. + * Non-integer values (e.g. 3.14) are printed as-is by the default stream. + * + * @param v Double value to format. + * @return String representation matching Python float str(). + */ +inline std::string format_double_python(double v) { + std::ostringstream oss; + oss.imbue(std::locale::classic()); // Ensure '.' decimal separator regardless of process locale. + oss << v; + std::string s = oss.str(); + if (s.find('.') == std::string::npos && s.find('e') == std::string::npos && + s.find('E') == std::string::npos) { + s += ".0"; + } + return s; +} + +} // namespace log_helpers + +#endif // INTERNALS_LOG_HELPERS_H_ diff --git a/feature_integration_tests/test_scenarios/cpp/src/internals/persistency/kvs_build_helpers.h b/feature_integration_tests/test_scenarios/cpp/src/internals/persistency/kvs_build_helpers.h index e0f65444aa8..e82e591597d 100644 --- a/feature_integration_tests/test_scenarios/cpp/src/internals/persistency/kvs_build_helpers.h +++ b/feature_integration_tests/test_scenarios/cpp/src/internals/persistency/kvs_build_helpers.h @@ -15,80 +15,22 @@ #define INTERNALS_PERSISTENCY_KVS_BUILD_HELPERS_H_ #include "kvs_parameters.h" +#include "internals/log_helpers.h" #include #include -#include -#include -#include #include -#include #include #include namespace kvs_build_helpers { -/** - * @brief Return the current UNIX timestamp as a decimal string (seconds). - * - * Used to populate the "timestamp" field in structured JSON log lines so that - * the C++ output matches the Rust tracing JSON shape expected by the FIT log - * filters. - * - * @return String containing the number of seconds since the UNIX epoch. - */ -inline std::string unix_seconds_string() { - const auto now = std::chrono::system_clock::now(); - const auto secs = - std::chrono::duration_cast(now.time_since_epoch()).count(); - return std::to_string(secs); -} - -/** - * @brief Emit a structured JSON INFO log line to stdout. - * - * Matches the Rust tracing JSON format expected by the FIT LogContainer so - * that Python test assertions can use find_log() uniformly for both Rust and - * C++ scenarios. - * - * Example output: - * @code - * {"timestamp":"1234567890","level":"INFO","fields":{"key":"my_key","value":42.0}, - * "target":"cpp_test_scenarios::scenarios::persistency::my_module","threadId":"ThreadId(1)"} - * @endcode - * - * @param fields JSON fragment for the "fields" object, e.g. @c "\"key\":\"x\",\"value\":1.0" - * @param target Module target string embedded in the log line. - */ -inline void log_info(const std::string& fields, const std::string& target) { - std::cout << "{\"timestamp\":\"" << unix_seconds_string() - << "\",\"level\":\"INFO\",\"fields\":{" << fields - << "},\"target\":\"" << target - << "\",\"threadId\":\"ThreadId(1)\"}\n"; -} - -/** - * @brief Format a double value to match Python's str(float) representation. - * - * For whole-number values (e.g. 42.0, 200.0) this appends ".0" so that the - * resulting string matches what Python's f-string interpolation produces. - * Non-integer values (e.g. 3.14) are printed as-is by the default stream. - * - * @param v Double value to format. - * @return String representation matching Python float str(). - */ -inline std::string format_double_python(double v) { - std::ostringstream oss; - oss.imbue(std::locale::classic()); // Ensure '.' decimal separator regardless of process locale. - oss << v; - std::string s = oss.str(); - if (s.find('.') == std::string::npos && s.find('e') == std::string::npos && - s.find('E') == std::string::npos) { - s += ".0"; - } - return s; -} +// Generic structured-log helpers live in the feature-neutral internals/log_helpers.h so the +// log line format (including json_escape of `target`) is defined once for all scenarios. +using log_helpers::format_double_python; +using log_helpers::log_info; +using log_helpers::unix_seconds_string; /** * @brief Convert an optional KvsDefaults mode to the boolean flag expected by KvsBuilder. diff --git a/feature_integration_tests/test_scenarios/cpp/src/scenarios/lifecycle/conditional_launching.cpp b/feature_integration_tests/test_scenarios/cpp/src/scenarios/lifecycle/conditional_launching.cpp new file mode 100644 index 00000000000..ef8450729d5 --- /dev/null +++ b/feature_integration_tests/test_scenarios/cpp/src/scenarios/lifecycle/conditional_launching.cpp @@ -0,0 +1,277 @@ +/******************************************************************************** + * Copyright (c) 2026 Contributors to the Eclipse Foundation + * + * See the NOTICE file(s) distributed with this work for additional + * information regarding copyright ownership. + * + * This program and the accompanying materials are made available under the + * terms of the Apache License Version 2.0 which is available at + * https://www.apache.org/licenses/LICENSE-2.0 + * + * SPDX-License-Identifier: Apache-2.0 + ********************************************************************************/ + +#include "conditional_launching.h" + +#include "internals/log_helpers.h" +#include "score/json/json_parser.h" + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +namespace { + +constexpr const char* kTarget = "cpp_test_scenarios::scenarios::lifecycle::conditional_launching"; + +void log_info(const std::string& message) { + log_helpers::log_info("\"message\":\"" + log_helpers::json_escape(message) + "\"", kTarget); +} + +bool path_condition_met(const std::string& path) { + std::error_code ec; + return std::filesystem::exists(path, ec) && !ec; +} + +bool env_condition_met(const std::string& name) { + return std::getenv(name.c_str()) != nullptr; +} + +// True if some process's /proc//comm (15-char truncated) or argv[0] basename equals +// `process_name`. Best effort: a process exiting mid-scan counts as "not found this pass". +bool process_condition_met(const std::string& process_name) { + // Iteration can still throw filesystem_error when a pid vanishes, despite the error_code overload. + try { + std::error_code ec; + for (const auto& entry : std::filesystem::directory_iterator( + "/proc", std::filesystem::directory_options::skip_permission_denied, ec)) { + const std::string pid = entry.path().filename().string(); + if (pid.empty() || + !std::all_of(pid.begin(), pid.end(), [](unsigned char c) { return std::isdigit(c); })) { + continue; + } + + std::ifstream comm(entry.path() / "comm"); + std::string comm_value; + if (comm && std::getline(comm, comm_value) && comm_value == process_name) { + return true; + } + + std::ifstream cmdline(entry.path() / "cmdline"); + std::stringstream cmdline_buffer; + cmdline_buffer << cmdline.rdbuf(); + const std::string argv0 = cmdline_buffer.str(); + if (!argv0.empty()) { + const auto argv0_end = argv0.find('\0'); + const std::string first_arg = argv0.substr(0, argv0_end); + // Basename, not suffix: "/usr/bin/oversleep" must not match "sleep". + const auto slash_pos = first_arg.find_last_of('/'); + const std::string basename = + slash_pos == std::string::npos ? first_arg : first_arg.substr(slash_pos + 1); + if (basename == process_name) { + return true; + } + } + } + } catch (const std::filesystem::filesystem_error&) { + return false; + } + return false; +} + +template +std::optional> parse_string_array_field(const std::string& input, + const std::string& field_name, + Converter convert) { + std::vector values; + + const score::json::JsonParser parser; + const auto root_any_res = parser.FromBuffer(input); + if (!root_any_res.has_value()) { + return std::nullopt; + } + + const auto root_object_res = root_any_res.value().As(); + if (!root_object_res.has_value()) { + return std::nullopt; + } + + const auto& root = root_object_res.value().get(); + const auto test_it = root.find("test"); + if (test_it == root.end()) { + return std::nullopt; + } + + const auto test_object_res = test_it->second.As(); + if (!test_object_res.has_value()) { + return std::nullopt; + } + + const auto& test = test_object_res.value().get(); + const auto field_it = test.find(field_name); + if (field_it == test.end()) { + return std::nullopt; + } + + const auto array_res = field_it->second.As(); + if (!array_res.has_value()) { + return std::nullopt; + } + + for (const auto& element : array_res.value().get()) { + const auto converted = convert(element); + if (!converted.has_value()) { + throw std::invalid_argument("Wait condition entries must be strings"); + } + values.push_back(*converted); + } + + return values; +} + +std::optional> parse_wait_conditions(const std::string& input) { + return parse_string_array_field(input, "wait_conditions", [](const score::json::Any& element) { + const auto value = element.As(); + if (!value.has_value()) { + return std::optional{}; + } + return std::optional{value.value()}; + }); +} + +// FIT stub (not launch_manager): validates `test.wait_conditions`, then polls each condition every +// `polling_interval_ms` until all are met or `timeout_ms` expires. A met condition stays latched, +// even if it later becomes false. Mirrors the Rust scenario, including its error messages. +class ConditionalLaunching : public Scenario { +public: + std::string name() const override { return "conditional_launching"; } + + void run(const std::string& input) const override { + const score::json::JsonParser parser; + const auto root_any_res = parser.FromBuffer(input); + if (!root_any_res.has_value()) { + throw std::invalid_argument("Failed to parse scenario input JSON"); + } + + uint64_t polling_interval = 50; + uint64_t timeout = 5000; + const auto wait_conditions_res = parse_wait_conditions(input); + + const auto root_object_res = root_any_res.value().As(); + if (root_object_res.has_value()) { + const auto& root = root_object_res.value().get(); + const auto test_it = root.find("test"); + if (test_it != root.end()) { + const auto test_object_res = test_it->second.As(); + if (test_object_res.has_value()) { + const auto& test = test_object_res.value().get(); + + const auto polling_it = test.find("polling_interval_ms"); + if (polling_it != test.end()) { + const auto polling_res = polling_it->second.As(); + if (polling_res.has_value()) { + polling_interval = polling_res.value(); + } + } + + const auto timeout_it = test.find("timeout_ms"); + if (timeout_it != test.end()) { + const auto timeout_res = timeout_it->second.As(); + if (timeout_res.has_value()) { + timeout = timeout_res.value(); + } + } + } + } + } + + // Same messages as the Rust scenario: "missing" and "empty" are distinct errors. + if (!wait_conditions_res.has_value()) { + throw std::runtime_error( + "Wait conditions were not provided: missing 'test.wait_conditions' in scenario input"); + } + const auto& wait_conditions = *wait_conditions_res; + if (wait_conditions.empty()) { + throw std::runtime_error( + "Wait conditions were not provided: empty 'test.wait_conditions' in scenario input"); + } + + log_info("Testing conditional launching"); + + for (const auto& condition : wait_conditions) { + if (condition.rfind("path:", 0) != 0U && condition.rfind("env:", 0) != 0U && + condition.rfind("process:", 0) != 0U) { + throw std::runtime_error("Unsupported wait condition prefix: " + condition); + } + } + + log_info("Polling interval: " + std::to_string(polling_interval) + "ms"); + log_info("Condition timeout: " + std::to_string(timeout) + "ms"); + + const auto deadline = std::chrono::steady_clock::now() + std::chrono::milliseconds(timeout); + std::vector satisfied(wait_conditions.size(), false); + + while (true) { + bool all_satisfied = true; + for (std::size_t i = 0; i < wait_conditions.size(); ++i) { + if (satisfied[i]) { + continue; + } + const auto& condition = wait_conditions[i]; + bool met = false; + if (condition.rfind("path:", 0) == 0U) { + met = path_condition_met(condition.substr(5)); + } else if (condition.rfind("env:", 0) == 0U) { + met = env_condition_met(condition.substr(4)); + } else { + met = process_condition_met(condition.substr(8)); + } + + if (met) { + satisfied[i] = true; + log_info("Condition satisfied: " + condition); + } else { + all_satisfied = false; + } + } + + if (all_satisfied) { + break; + } + + if (std::chrono::steady_clock::now() >= deadline) { + std::string unmet; + for (std::size_t i = 0; i < wait_conditions.size(); ++i) { + if (!satisfied[i]) { + if (!unmet.empty()) { + unmet += ", "; + } + unmet += wait_conditions[i]; + } + } + throw std::runtime_error("Timed out after " + std::to_string(timeout) + + "ms waiting for condition(s): " + unmet); + } + + std::this_thread::sleep_for(std::chrono::milliseconds(polling_interval)); + } + + log_info("All dependencies satisfied"); + } +}; + +} // namespace + +Scenario::Ptr make_conditional_launching_scenario() { + return std::make_shared(); +} diff --git a/feature_integration_tests/test_scenarios/cpp/src/scenarios/lifecycle/conditional_launching.h b/feature_integration_tests/test_scenarios/cpp/src/scenarios/lifecycle/conditional_launching.h new file mode 100644 index 00000000000..95a4a6a1797 --- /dev/null +++ b/feature_integration_tests/test_scenarios/cpp/src/scenarios/lifecycle/conditional_launching.h @@ -0,0 +1,17 @@ +/******************************************************************************** + * Copyright (c) 2026 Contributors to the Eclipse Foundation + * + * See the NOTICE file(s) distributed with this work for additional + * information regarding copyright ownership. + * + * This program and the accompanying materials are made available under the + * terms of the Apache License Version 2.0 which is available at + * https://www.apache.org/licenses/LICENSE-2.0 + * + * SPDX-License-Identifier: Apache-2.0 + ********************************************************************************/ +#pragma once + +#include + +Scenario::Ptr make_conditional_launching_scenario(); diff --git a/feature_integration_tests/test_scenarios/cpp/src/scenarios/mod.cpp b/feature_integration_tests/test_scenarios/cpp/src/scenarios/mod.cpp index 83a32e5af8e..67bac4d8a2e 100644 --- a/feature_integration_tests/test_scenarios/cpp/src/scenarios/mod.cpp +++ b/feature_integration_tests/test_scenarios/cpp/src/scenarios/mod.cpp @@ -13,6 +13,8 @@ #include +#include "scenarios/lifecycle/conditional_launching.h" + #include Scenario::Ptr make_multiple_kvs_per_app_scenario(); @@ -38,9 +40,18 @@ ScenarioGroup::Ptr persistency_scenario_group() { std::vector{supported_datatypes_group(), default_values_group()}); } +ScenarioGroup::Ptr lifecycle_scenario_group() { + return std::make_shared( + "lifecycle", + std::vector{ + make_conditional_launching_scenario(), + }, + std::vector{}); +} + ScenarioGroup::Ptr root_scenario_group() { return std::make_shared( "root", std::vector{}, - std::vector{persistency_scenario_group()}); + std::vector{persistency_scenario_group(), lifecycle_scenario_group()}); } diff --git a/feature_integration_tests/test_scenarios/rust/src/main.rs b/feature_integration_tests/test_scenarios/rust/src/main.rs index 024b09a2555..aedd088a3fc 100644 --- a/feature_integration_tests/test_scenarios/rust/src/main.rs +++ b/feature_integration_tests/test_scenarios/rust/src/main.rs @@ -22,6 +22,7 @@ use std::time::{SystemTime, UNIX_EPOCH}; use tracing::Level; use tracing_subscriber::fmt::time::FormatTime; use tracing_subscriber::FmtSubscriber; + struct NumericUnixTime; impl FormatTime for NumericUnixTime { diff --git a/feature_integration_tests/test_scenarios/rust/src/scenarios/lifecycle/conditional_launching.rs b/feature_integration_tests/test_scenarios/rust/src/scenarios/lifecycle/conditional_launching.rs new file mode 100644 index 00000000000..f3fc0d8fe2b --- /dev/null +++ b/feature_integration_tests/test_scenarios/rust/src/scenarios/lifecycle/conditional_launching.rs @@ -0,0 +1,159 @@ +// ******************************************************************************* +// Copyright (c) 2026 Contributors to the Eclipse Foundation +// +// See the NOTICE file(s) distributed with this work for additional +// information regarding copyright ownership. +// +// This program and the accompanying materials are made available under the +// terms of the Apache License Version 2.0 which is available at +// +// +// SPDX-License-Identifier: Apache-2.0 +// ******************************************************************************* + +use serde_json::Value; +use std::fs; +use std::path::Path; +use std::time::{Duration, Instant}; +use test_scenarios_rust::scenario::Scenario; +use tracing::info; + +/// FIT stub (not launch_manager): validates `test.wait_conditions`, then polls each condition every +/// `polling_interval_ms` until all are met or `timeout_ms` expires. A met condition stays latched, +/// even if it later becomes false. +pub struct ConditionalLaunching; + +fn path_condition_met(path: &str) -> bool { + Path::new(path).exists() +} + +fn env_condition_met(name: &str) -> bool { + std::env::var_os(name).is_some() +} + +/// True if some process's /proc//comm (15-char truncated) or argv[0] basename equals +/// `process_name`. Best effort: unreadable entries are skipped. +fn process_condition_met(process_name: &str) -> bool { + let Ok(entries) = fs::read_dir("/proc") else { + return false; + }; + + for entry in entries.flatten() { + let pid = entry.file_name(); + let Some(pid) = pid.to_str() else { continue }; + if !pid.chars().all(|c| c.is_ascii_digit()) { + continue; + } + + if let Ok(comm) = fs::read_to_string(entry.path().join("comm")) { + if comm.trim_end() == process_name { + return true; + } + } + + if let Ok(cmdline) = fs::read(entry.path().join("cmdline")) { + let argv0 = cmdline.split(|&b| b == 0).next().unwrap_or(&[]); + if let Ok(argv0) = std::str::from_utf8(argv0) { + // Basename, not suffix: "/usr/bin/oversleep" must not match "sleep". + let basename = argv0.rsplit('/').next().unwrap_or(argv0); + if basename == process_name { + return true; + } + } + } + } + false +} + +impl Scenario for ConditionalLaunching { + fn name(&self) -> &str { + "conditional_launching" + } + + fn run(&self, input: &str) -> Result<(), String> { + let value: Value = serde_json::from_str(input).map_err(|error| format!("Parse error: {error}"))?; + let test = value + .get("test") + .ok_or_else(|| "Missing 'test' field in scenario input".to_string())?; + + let polling_interval = test.get("polling_interval_ms").and_then(Value::as_u64).unwrap_or(50); + let timeout = test.get("timeout_ms").and_then(Value::as_u64).unwrap_or(5000); + let conditions = test.get("wait_conditions").and_then(Value::as_array).ok_or_else(|| { + "Wait conditions were not provided: missing 'test.wait_conditions' in scenario input".to_string() + })?; + + if conditions.is_empty() { + return Err( + "Wait conditions were not provided: empty 'test.wait_conditions' in scenario input".to_string(), + ); + } + + info!("Testing conditional launching"); + + let conditions: Vec<&str> = conditions + .iter() + .map(|condition| { + condition + .as_str() + .ok_or_else(|| "Wait condition entries must be strings".to_string()) + }) + .collect::>()?; + + for condition in &conditions { + if !condition.starts_with("path:") && !condition.starts_with("env:") && !condition.starts_with("process:") { + return Err(format!("Unsupported wait condition prefix: {condition}")); + } + } + + info!("Polling interval: {polling_interval}ms"); + info!("Condition timeout: {timeout}ms"); + + let deadline = Instant::now() + Duration::from_millis(timeout); + let mut satisfied = vec![false; conditions.len()]; + + loop { + let mut all_satisfied = true; + for (index, condition) in conditions.iter().enumerate() { + if satisfied[index] { + continue; + } + + let met = if let Some(path) = condition.strip_prefix("path:") { + path_condition_met(path) + } else if let Some(name) = condition.strip_prefix("env:") { + env_condition_met(name) + } else { + process_condition_met(condition.strip_prefix("process:").expect("checked above")) + }; + + if met { + satisfied[index] = true; + info!("Condition satisfied: {condition}"); + } else { + all_satisfied = false; + } + } + + if all_satisfied { + break; + } + + if Instant::now() >= deadline { + let unmet = conditions + .iter() + .zip(&satisfied) + .filter(|(_, met)| !**met) + .map(|(condition, _)| *condition) + .collect::>() + .join(", "); + return Err(format!("Timed out after {timeout}ms waiting for condition(s): {unmet}")); + } + + std::thread::sleep(Duration::from_millis(polling_interval)); + } + + info!("All dependencies satisfied"); + + Ok(()) + } +} diff --git a/feature_integration_tests/test_scenarios/rust/src/scenarios/lifecycle/mod.rs b/feature_integration_tests/test_scenarios/rust/src/scenarios/lifecycle/mod.rs new file mode 100644 index 00000000000..2c180f72b89 --- /dev/null +++ b/feature_integration_tests/test_scenarios/rust/src/scenarios/lifecycle/mod.rs @@ -0,0 +1,25 @@ +// ******************************************************************************* +// Copyright (c) 2026 Contributors to the Eclipse Foundation +// +// See the NOTICE file(s) distributed with this work for additional +// information regarding copyright ownership. +// +// This program and the accompanying materials are made available under the +// terms of the Apache License Version 2.0 which is available at +// +// +// SPDX-License-Identifier: Apache-2.0 +// ******************************************************************************* + +mod conditional_launching; + +use conditional_launching::ConditionalLaunching; +use test_scenarios_rust::scenario::{ScenarioGroup, ScenarioGroupImpl}; + +pub fn lifecycle_group() -> Box { + Box::new(ScenarioGroupImpl::new( + "lifecycle", + vec![Box::new(ConditionalLaunching)], + vec![], + )) +} diff --git a/feature_integration_tests/test_scenarios/rust/src/scenarios/mod.rs b/feature_integration_tests/test_scenarios/rust/src/scenarios/mod.rs index 5c7013138f6..2ab4d9c956c 100644 --- a/feature_integration_tests/test_scenarios/rust/src/scenarios/mod.rs +++ b/feature_integration_tests/test_scenarios/rust/src/scenarios/mod.rs @@ -12,10 +12,16 @@ // ******************************************************************************* use test_scenarios_rust::scenario::{ScenarioGroup, ScenarioGroupImpl}; +mod lifecycle; mod persistency; +use lifecycle::lifecycle_group; use persistency::persistency_group; pub fn root_scenario_group() -> Box { - Box::new(ScenarioGroupImpl::new("root", vec![], vec![persistency_group()])) + Box::new(ScenarioGroupImpl::new( + "root", + vec![], + vec![lifecycle_group(), persistency_group()], + )) }