The module lifecycle stage: Generally available version

The module has requirements for installation

Overview

Network equipment — switches, routers, firewalls, UPS units, storage systems, printers — is polled over SNMP by opagent-snmp, a subagent of the monitoring agent. It is a separate process that the agent supervisor downloads, starts and keeps up to date next to the agent on the same node, so the SNMP polling can be updated and restarted without touching the host monitoring itself.

Polling runs from the node, not from the cluster: the node needs network access to the devices, the cluster does not.

The subagent is off by default. It starts polling when both of the following are done:

  1. The subagent is assigned to the hosts in the administrator web interface.
  2. A configuration file listing the devices is placed on the node.

Without its configuration file the process runs idle and polls nothing.

Devices are polled once a minute. All samples of one cycle get a single timestamp aligned to the minute grid, so rate() over the counters does not break when the duration of a poll varies.

Enabling

Assigning the subagent

The subagent and its version are assigned in Admin → Monitoring agents → Subagents:

  • Default set — what every host without an override receives. Add a row, choose the subagent (opagent-snmp), choose its version and, if needed, the CPU (mCPU) and memory (bytes) limits. The versions offered are the ones published into the installation’s storage by the module, so a newly released subagent version appears here after the module is updated.
  • Pinned version — versions and limits pinned to a specific project or host. A pin takes priority over the default set.
  • State — what the supervisors report back: the running and the desired version of the subagent per host.

For the general description of subagents and the way a version is assigned, see Web Interface → Subagents.

Placing the configuration on the node

The subagent reads one YAML file from the agent’s plugin configuration directory:

  • the directory is plugins_config_dir from the agent’s config.yaml, and /usr/local/<product>/etc/config.d/ when the agent does not set it (<product> is the monitoring.agent.productName value, okagent by default);
  • the file is the first one matching *.yaml, then *.yml, whose top-level key plugin equals snmp. The file name itself does not matter;
  • the device list lives under the key config.

The directory is rescanned once a minute, so adding, changing or removing the file takes effect without restarting anything.

A minimal configuration — one device, one authentication profile:

plugin: snmp              # the subagent finds its file by this key
config:
  targets:
    - host: 192.168.1.1
  auths:
    default:              # with a single profile a target need not name it
      community: public
      version: 2

Configuration

Top-level fields

Field Type Default Description
targets list of targets — devices to poll
auths map, name → profile — authentication profiles
default_modules list of strings if_mib, system, host_resources, lldp, bridge, qbridge modules for targets that do not name their own; an empty list means the same defaults
custom_mibs_file path — a file in the same format as the built-in module set; its top-level modules and auths are laid over the built-in ones by name
limit integer 10000 cap on the number of series one polling cycle ships; the tail above the cap is dropped
scrape_timeout duration string, e.g. 45s not set deadline of a target’s main polling phase; unset and 0s mean no deadline of its own
discovery section — scanning subnets for devices

A key that is not in the table is ignored silently.

Target

Field Type Default Description
host string — device address
auth profile name the only profile in auths with more than one profile the name is required
modules list of strings default_modules an explicit list turns auto-classification off for this target
vlans list of integers empty — VLANs are discovered automatically applies only where the classifier assigned per-VLAN modules, that is, not to a target with an explicit modules

Authentication profile

Field Type Default Description
community string public, while version is below 3 SNMPv1 and v2c
version integer 2 1 — v1, 3 — USM (SNMPv3), anything else — v2c
security_level string noAuthNoPriv noAuthNoPriv, authNoPriv, authPriv; case does not matter
username string — USM user name, v3
password string — authentication passphrase, v3
auth_protocol string MD5 MD5, SHA, SHA224, SHA256, SHA384, SHA512
priv_protocol string DES DES, AES, AES192, AES192C, AES256, AES256C
priv_password string — privacy passphrase, v3
context_name string — SNMP context of the main polling phase

Discovery

Discovery runs once, on the first polling cycle. The addresses it finds become targets and go through auto-classification; hosts already listed in targets are skipped. The probes use the profile named discovery from auths, and when there is no such profile, the fields of this section.

Field Type Default Description
subnets list of IPv4 CIDRs — what to scan. An unparseable subnet and IPv6 are reported in the log and skipped; the rest is still scanned. For a prefix shorter than /31 the network and broadcast addresses are skipped
community string public community of the probes
version integer v2c 1 — v1, anything else — v2c
timeout duration string 2s timeout of one probe
concurrency integer 32 how many addresses are probed at once

Validation

A configuration that would poll nothing, or poll it with credentials nobody chose, is rejected:

  • at least one target or a non-empty discovery.subnets is required;
  • a non-empty auths is required when discovery.subnets is empty;
  • every target needs a non-empty host;
  • targets[].auth must name a key of auths; a target that names none is only accepted when auths holds exactly one profile;
  • scrape_timeout must parse, and neither it nor limit may be negative.

The SNMP version is not validated: an unknown one is read as v2c.

Example with SNMPv3, discovery and explicit modules

plugin: snmp
config:
  scrape_timeout: 45s     # deadline of the main phase; without the key there is none
  limit: 20000            # series cap per cycle, 10000 by default
  default_modules: [if_mib, system, lldp]
  auths:
    v3:
      version: 3
      username: monitor
      security_level: authPriv
      auth_protocol: SHA256
      password: <auth passphrase>
      priv_protocol: AES256
      priv_password: <priv passphrase>
    ro: {community: public, version: 2}
  targets:
    - host: 10.0.0.1
      auth: v3            # with more than one profile the name is required
    - host: 10.0.0.2
      auth: ro
      modules: [if_mib, system]   # an explicit list: no auto-classification
  discovery:
    subnets: [10.0.1.0/24]
    community: public     # the "discovery" profile is assembled from here

Device auto-classification

A target that does not name its own modules is asked for sysObjectID on the first cycle, and the modules of the matching vendor are added to the default ones. There are 47 rules over enterprise OID prefixes, covering Cisco (including WLC), Huawei, Juniper, Arista, MikroTik, Eltex, QTECH, Extreme, Allied Telesis, D-Link, Netgear, TP-Link, Zyxel, Ubiquiti, Ruckus, NEC IX, HPE and Aruba, Dell (both servers and Dell Networking), Fortinet, Palo Alto, Check Point, Sophos, F5, Citrix, Kemp, Brocade, NetApp, QNAP, Synology, APC and other UPS vendors, and printers.

A device that does not answer the probe, or whose prefix matches no rule, keeps the default modules. A target with an explicit modules is not classified at all.

Metrics

The data of the SNMP subagent is a separate series stream with its own names and labels; it is not the netdev plugin of the agent and does not reuse its metric names. For the reference, see Agent metrics → SNMP subagent.

Labels

Label Value
job snmp
instance host name of the polling node: source_hostname from the agent configuration, otherwise the node’s own host name
target address of the polled device
vlan VLAN ID, only on the series of the per-VLAN phase

Device metric names are the MIB object names (ifHCInOctets, sysUpTime, lldpRemSysName, and so on) and are kept as they are; text values arrive as info metrics (ifType_info{ifType="…"} 1).

Subagent’s own metrics

Metric Meaning
snmp_scrape_success{target} 1 — the device answered and the cycle collected data, 0 — it did not; the text of the error goes to the log
snmp_scrape_duration_seconds{target} how long polling the device took, both phases
snmp_config_status{state="ok|missing|invalid"} state of the subagent’s configuration file, shipped on every cycle
snmp_build_info{version} the running version, written on every tick

There is no separate heartbeat series: liveness of the subagent is time() - timestamp(snmp_build_info).

Infrastructure map and device list

Polled devices are more than a metric stream: the platform builds the project’s device list and its infrastructure map out of them. Both pages live in the project sidebar, Infrastructure → Devices and Infrastructure → Map; they appear once the administrator has switched the infrastructure section on in the platform settings and the SNMP infrastructure UI on in the project settings. Both are rebuilt from the stored series every ten minutes.

  • Devices. A device enters the list as soon as its sysName arrives, one entry per polled address. What the entry says — vendor, category (switch, router, firewall, UPS and so on), location, contact — is read from the system metrics of the device itself: sysObjectID, sysDescr, sysLocation, sysContact; the ENTITY-MIB entities tell a switch stack from a standalone one. The device page adds its ports — status, speed, traffic, errors, discards — and the links of that device.
  • Map. Links between devices come from LLDP (lldpRemSysName, lldpRemChassisId, lldpRemPortId) and, where LLDP says nothing, from CDP (cdpCacheDeviceId). A link from a switch port to a server comes from the bridge table (dot1dTpFdbPort): the MAC learned on the port is matched against the MAC addresses the monitoring agent reports for the server. Only a port with exactly one server behind it becomes such a link, so uplinks and inter-switch trunks do not produce false ones.
  • An LLDP neighbour that is not polled still gets an entry of its own, under “Discovered devices”: it is on the map as a neighbour, but has no metrics. Add its address to targets and it becomes a polled device.

A device whose sysName has not arrived for an hour is marked stale, and one not seen for thirty days is removed; links behave the same way.

Dashboards

Ready-made SNMP dashboards come with the module:

  • SNMP Devices — the fleet: all polled devices, their availability and load.
  • SNMP Device Detailed — one device: system information, interfaces, and their traffic and errors.

Troubleshooting

Start from snmp_config_status: it is shipped on every cycle, even when nothing is being polled.

Symptom What it looks like What to do
The configuration was not picked up snmp_config_status{state="missing"} there is no file with plugin: snmp in the directory. This is also the normal state of a subagent nobody has configured yet. Check the path and the plugin key; the directory is rescanned once a minute, no restart is needed. A duplicate key inside one YAML block makes the file unreadable and looks the same way from the outside
The configuration is broken snmp_config_status{state="invalid"} the file is there but a poller could not be built from it. The log names the file and the reason. The previously applied configuration keeps working, so a typo does not stop the polling that is already running
A device does not answer snmp_scrape_success{target} 0 wrong community or USM parameters, no network access to the device, or the device is down. The text of the error is in the log
Polling does not fit into the interval snmp_scrape_duration_seconds{target} approaching or above 60 the main phase is all-or-nothing: when the configured scrape_timeout fires, the whole cycle’s data for that device is dropped and snmp_scrape_success becomes 0. By default there is no deadline at all — a slow switch with a large interface table returns complete data even when it does not fit into a tick. Set scrape_timeout only when the device is known to answer faster
Nothing at all arrives from the node snmp_build_info stops updating the subagent is not running. Check its state in Admin → Monitoring agents → Subagents, section “State”

Requirements

  • Network access from the polling node to the equipment over SNMP: UDP, port 161. The subagent’s own probes — the sysObjectID classification probe, VLAN discovery and the subnet scan — always go to port 161.
  • Credentials for the equipment: a community string for SNMPv1 and v2c, or a USM user for SNMPv3.
  • Polling runs from the node the agent is installed on, so it is the node’s network access that matters, not the cluster’s.
  • For per-VLAN polling of bridge tables on Cisco equipment, an SNMPv3 user with access to the vlan- SNMP contexts.