Observability Engineer

 

Description:

We are seeking an experienced IoT / Observability Engineer responsible for monitoring, managing, automating, and optimizing enterprise infrastructure, applications, IoT platforms, and enterprise telemetry ecosystems. The ideal candidate will possess strong expertise in observability platforms, monitoring technologies, automation frameworks, and operational support processes to ensure high availability, reliability, and performance of business-critical systems.
The engineer will support enterprise monitoring operations, FOAK services, ELT platforms, incident management, and automation initiatives while collaborating with Infrastructure, Cloud, Network, Application, and Service Delivery teams.
We want all new Associates to succeed in their roles at Ensono. That's why we've outlined the job requirements below. To be considered for this role, it's important that you meet all Required Qualifications. If you do not meet all of the Preferred Qualifications, we still encourage you to apply. 
Key Responsibilities

Monitoring & Observability

Design, implement, and maintain enterprise monitoring and observability solutions.
Develop and maintain dashboards, alerts, and visualizations using Grafana.
Monitor infrastructure, applications, middleware, and IoT services using IBM Instana, Grafana, SolarWinds, and related observability tools.
Configure and manage data collection using Telegraf, Prometheus, and monitoring agents.
Analyze metrics, logs, traces, events, and telemetry data to identify performance bottlenecks and service degradation.
Support SLO, SLA, and operational health monitoring initiatives.
Perform Root Cause Analysis (RCA) and troubleshooting for infrastructure and application issues.
Foak & Enterprise Logging/Telemetry

Support onboarding, monitoring, and operational management of FOAK (First Office Application Kit) services and enterprise applications.
Configure, validate, and troubleshoot Enterprise Logging & Telemetry (ELT) integrations across infrastructure, middleware, applications, and cloud platforms.
Monitor telemetry pipelines, log ingestion, event correlation, and data quality to ensure complete observability coverage.
Collaborate with engineering teams to improve telemetry standards, monitoring effectiveness, and proactive incident detection through ELT and observability frameworks.
Support FOAK application integrations with Grafana, Instana, Prometheus, and enterprise monitoring platforms.
Infrastructure & Platform Monitoring

Monitor and support:
Linux Servers
Windows Servers
VMware Infrastructure
Citrix VDI Platforms
DNS Services
Proxy Services
Middleware Platforms
Integration Services
Enterprise Applications
IoT Platforms
Additional Responsibilities:
Investigate performance issues, recurring alerts, and infrastructure anomalies.
Validate monitoring platform health and monitoring coverage.
Monitor capacity, availability, CPU, memory, storage, and service health metrics.
Support platform upgrades, maintenance, and operational readiness reviews.
Database & Data Management

Configure and maintain InfluxDB time-series databases.
Manage data retention policies, performance tuning, and capacity planning.
Develop operational dashboards and reports for infrastructure and application performance insights.
Support telemetry data ingestion, storage optimization, and historical trend analysis.
Event & Incident Management

Monitor operational alerts, events, notifications, and incidents from enterprise monitoring platforms.
Acknowledge, investigate, troubleshoot, and resolve assigned incidents.
Coordinate with Infrastructure, Network, Cloud, Security, Application, and Service Delivery teams during incident resolution.
Participate in major incident bridges, DR exercises, and 24x7 operations support activities.
Follow escalation procedures, SOPs, operational runbooks, and ITIL processes.
Support Problem Management activities and contribute to RCA documentation.
Instana & APM Operations

Administer and support IBM Instana monitoring environments.
Monitor application, API, middleware, and microservices performance using Instana.
Validate Instana agent health following server patching and maintenance activities.
Support application onboarding and APM configuration standards.
Configure alerts, baselines, and performance thresholds.
Raise and track incidents related to Instana platform availability and performance.
Automation & Scripting

Develop automation solutions using:
Python
PowerShell
Shell Scripting (Bash)
VBScript
 Responsibilities:
Automate operational tasks, monitoring deployments, and remediation workflows.
Build reusable automation tools to improve operational efficiency.
Integrate monitoring platforms with enterprise automation frameworks.
Support webhook-based automation and event-driven operational workflows.
Configuration Management & Infrastructure Automation

Implement Infrastructure as Code (IaC) and automation using:
Ansible
Puppet
Responsibilities:
Automate server provisioning and configuration management.
Automate monitoring agent deployment and onboarding.
Maintain automation playbooks and deployment pipelines.
Improve operational consistency and reduce manual efforts across environments.
Knowledge Management

Maintain SOPs, runbooks, monitoring procedures, and escalation matrices.
Participate in KT sessions, service onboarding, and operational readiness reviews.
Support service transition, migration, and continuous improvement initiatives.
Maintain observability standards and monitoring documentation.
Required Skills

Monitoring & Observability

Grafana
IBM Instana
SolarWinds
Telegraf
Prometheus
InfluxDB
OpenTelemetry
Grafana Alloy
APM Monitoring
Event Management
Alert Management
Observability Concepts
SLO/SLA Monitoring
FOAK & Enterprise Telemetry

FOAK (First Office Application Kit) Support
Enterprise Logging & Telemetry (ELT)
Log Aggregation & Correlation
Telemetry Data Analysis
Event Correlation
Application Onboarding
Monitoring Standards & Observability Frameworks
Infrastructure

VMware
Linux Administration
Windows Server
Citrix VDI
DNS Services
Proxy Services
Middleware Technologies
Infrastructure Performance Monitoring
Scripting & Programming

Python
PowerShell
Shell Scripting (Linux)
VBScript
Automation Tools

Ansible
Puppet
Webhooks
Infrastructure as Code (IaC)
ITSM & Operations

ServiceNow
Incident Management
Problem Management
Change Management
ITIL Framework
Major Incident Management
Nice to Have

Docker
Kubernetes
AWS
Microsoft Azure
Google Cloud Platform (GCP)
Jenkins
GitHub Actions
GitLab CI/CD
REST APIs
Microservices Monitoring
DevOps & SRE Practices
Preferred Experience

7 to 10+ years of experience in Monitoring, Observability, Infrastructure Operations, SRE, or Platform Engineering.
Experience supporting large-scale enterprise environments and 24x7 operations.
Hands-on experience with Grafana, Instana, SolarWinds, Telegraf, Prometheus, and InfluxDB.
Experience supporting VMware, Citrix, middleware, enterprise applications, and cloud monitoring platforms.
Experience working with FOAK applications and Enterprise Logging & Telemetry (ELT) platforms.
Strong troubleshooting, RCA, incident management, and operational support skills.
Experience with automation frameworks and Infrastructure as Code (Ansible preferred).
Experience integrating observability platforms with enterprise automation solutions.
Key Competencies

Observability & Monitoring
Infrastructure Automation
Enterprise Telemetry Management
FOAK Application Support
Root Cause Analysis
Problem Solving
Performance Optimization
Service Reliability Engineering (SRE)
Cross-Functional Collaboration
Operational Excellence
Continuous Improvement

Organization Ensono
Industry Engineering Jobs
Occupational Category Observability Engineer
Job Location Illinois,USA
Shift Type Morning
Job Type Full Time
Gender No Preference
Career Level Experienced Professional
Experience 7 Years
Posted at 2026-09-17 3:12 pm
Expires on 2026-11-01