Skip to main content
Let's talk
Artificial Intelligence · AgentOps

Agents in production, under control

An agent is not delivered and forgotten. We operate your agents with continuous evals, step-level observability, drift detection, auditing and human control, so they keep doing the right thing when the data, the processes or the model change.

  • Traces for every decision and every tool used
  • Regression evals before any prompt or model change
  • Agent red teaming against injection and tool abuse

Design principles

  • Every run traced
  • No model or prompt change without evals
  • Automatic alerts on behavior and cost
  • One audit record per action

The hard part starts after launch

Models get new versions, data changes and processes change. An agent that worked well can degrade without anyone noticing.

AgentOps is the discipline of running agents as critical systems: measuring, detecting, auditing and correcting, with people at the right checkpoint.

Read: AgentOps, running agents in production →

What we operate

Agent observability

Traces for every step (inputs, decisions, tools, costs, latency), built on open standards such as OpenTelemetry.

Continuous evals

Custom test suites that run on every change and periodically in production.

How we evaluate agents →
Drift detection

Alerts when the agent's behavior, quality or cost drifts from its baseline.

Guardrails and human control

Autonomy levels per action type, approvals and operational limits.

Guardrails and human control →
Agent red teaming

Offensive testing against instruction injection, data leakage and tool abuse, with our ethical hacking team.

Cybersecurity and ethical hacking →
Governance and audit

An inventory of agents, owners, versions, permissions and evidence for internal audit, using frameworks such as the NIST AI RMF and ISO/IEC 42001 as a reference.

How we run it (indicative timelines)

011–2 weeks

Inventory and baseline

Which agents exist, who owns them, which versions are running and how they behave today.

021–2 weeks

Instrumentation and dashboards

Traces, quality and cost metrics, and a dashboard per agent.

032 weeks

Evals and alerts

Custom test suites and alerts on quality, drift and cost.

041–2 weeks

Initial red teaming

Agent-specific offensive testing and a remediation plan.

05Ongoing

Operation with periodic reports

Reports on quality, incidents, costs and improvements applied.

The stack we work with

Observability

OpenTelemetryLangfuseLangSmithGrafana

Cloud platforms

Amazon CloudWatchAzure Monitor

Quality and control

Custom evalsMCPRed teaming

Before building agents for others, we use them ourselves

What we offer is what we use every day.

Niveleads: agents in our sales operation

In Niveleads, our commercial platform, agents prospect, build campaigns, send automated emails and review how each lead engages. They also make AI voice calls, run technology watch, support growth and marketing, and read signals to give the team insights.

Agent-assisted engineering with independent QA

We build software with agents that write, test and review code, but we do not leave everything to AI. Deliveries go through an independent QA agent and verification standards defined by our team. Whoever builds is not whoever approves.

A website built for agents

nivelics.com is designed to be read by agents and LLMs as well: a dynamically generated llms.txt, schema.org structured data and content written with LLMs in mind.

Reference architectures

Illustrative scenarios of how we build. They do not describe clients or results.

Operations agent: from alert to deployment

An organization with many systems and an operations team flooded with alerts.

  1. Alert
  2. Triage agent
  3. Changes, logs and runbooks via MCP
  4. Verifier
  5. Execute and notify, or request approval
  6. Audit log

The agent correlates each alert with recent deployments and changes, queries logs and runbooks, classifies severity and proposes an action. Reversible actions pre-approved in the runbook, such as restarting a service, run once the verifier signs off, and the team is notified. For irreversible ones it prepares the plan and waits for human approval.

Harness: Read-only tools by default · verifier checks against the runbook · traces for every decision · evals built from past incidents

Multi-system orchestration of a business process

A process that crosses several systems and teams, with rework caused by inconsistent data.

  1. Request
  2. Orchestrator
  3. Integration and execution agents
  4. Verifier
  5. Close, or exception to a person

An orchestrator receives the request. An integration agent validates data across the ERP, the document system and the database. An execution agent records the changes, and an independent verifier confirms consistency before closing. Exceptions reach a person with all the context already assembled.

Harness: One least-privilege MCP server per system · idempotent transactions with compensation · model routing · AgentOps dashboard

Frequently asked questions

Yes. We start with an inventory and a baseline, instrument whatever is missing and define evals before taking over operations.

It is the gradual deviation of its behavior from what is expected, caused by changes in the data, the processes or the model version. It is detected by comparing against a baseline and with periodic evals.

Every run: what the agent received, what it decided, which tools it used with which parameters and permissions, and what result it got.

Through live dashboards and a periodic report on quality, incidents, costs and improvements applied.

Yes, with agent-specific red teaming: instruction injection, tool abuse and data leakage.

Want to audit your agents' security? See ethical hacking →

How many agents do you have in production, and who is watching them?

We start with an inventory and a baseline: you will know what your agents do, what they cost and where the risks are.

Talk about AgentOps →Reply in under 24 hours · Under a confidentiality agreement