We are looking for a NOC Engineer responsible for monitoring, managing, troubleshooting, and maintaining the availability and performance of our network, servers, telephony, and production infrastructure.
The candidate will work closely with the DevOps, SIP/VoIP, CTI, Development, and QA teams to identify incidents, troubleshoot infrastructure and connectivity issues, and ensure timely resolution of production-impacting problems.
Key Responsibilities
- Monitor production infrastructure, servers, networks, applications, and telephony services 24×7 as per the defined shift schedule.
- Monitor CPU, memory, disk, network utilization, latency, packet loss, service availability, and system health.
- Monitor SIP/VoIP infrastructure, SIP trunks, channels, call connectivity, registration, call failures, and related telephony services.
- Troubleshoot network connectivity issues involving TCP/IP, DNS, HTTP/HTTPS, VPN, firewalls, routing, and ports.
- Monitor AWS infrastructure, including EC2, ECS, RDS, Load Balancers, CloudWatch, and other relevant services.
- Identify production incidents proactively through monitoring and alerting systems.
- Perform first-level troubleshooting and determine whether an issue is related to network, server, application, database, SIP/telephony, or third-party services.
- Escalate incidents to the appropriate DevOps, Development, SIP/VoIP, Database, or third-party team with complete technical details.
- Maintain incident logs, escalation records, and resolution documentation.
- Perform health checks before and after production deployments.
- Monitor service uptime, system performance, and infrastructure capacity.
- Coordinate with telecom/SIP providers for issues related to SIP trunks, channels, numbers, call routing, connectivity, and call quality.
- Analyze logs and network-level information to identify the root cause of incidents.
- Support troubleshooting using tools such as Ping, Traceroute, Telnet/Netcat, nslookup/dig, curl, SSH, TCPDump, Wireshark, and similar utilities.
- Monitor alerts and ensure incidents are acknowledged and resolved within defined SLA/response timelines.
- Prepare daily/weekly operational reports covering incidents, outages, recurring issues, and system health.
- Maintain proper documentation for infrastructure, monitoring, escalation matrices, and standard operating procedures.
- Participate in incident management, problem management, and post-incident RCA activities.
- Support infrastructure capacity planning by identifying trends in resource utilization and service demand.
Required Technical Skills
- Networking
- Strong understanding of TCP/IP, UDP, DNS, DHCP, HTTP/HTTPS, VPN, NAT, routing, and firewalls.
- Understanding of network troubleshooting and connectivity diagnostics.
- Knowledge of LAN/WAN and network monitoring.
- Basic understanding of load balancing and reverse proxies.
- Linux & Servers
- Good working knowledge of Linux/Ubuntu/CentOS environments.
- Comfortable with SSH and Linux command-line troubleshooting.
- Understanding of processes, services, disk, memory, CPU, and network troubleshooting.
- Basic shell scripting knowledge is an advantage.
- AWS / Cloud
- Working knowledge of AWS infrastructure.
- Familiarity with EC2, ECS, RDS, CloudWatch, IAM, VPC, Security Groups, Load Balancers, and related services.
- Ability to troubleshoot basic AWS connectivity and infrastructure issues.
SIP / VoIP / Telephony
- Experience with SIP/VoIP or telecom environments will be highly preferred.
- The candidate should have an understanding of
- SIP signaling and call flow
- SIP registration
- INVITE / BYE / ACK / CANCEL / OPTIONS
- SIP response codes
- RTP and media flow
- SIP trunking
- Concurrent channels
- Call routing
- Caller ID
- Call failures and disconnects
- Basic VoIP call-quality parameters such as latency, jitter, and packet loss
Monitoring & Troubleshooting
- Experience with monitoring/logging tools such as -
- Grafana
- Prometheus
- AWS CloudWatch
- ELK / OpenSearch
- Zabbix
- Nagios
- Datadog
- Other infrastructure/network monitoring platforms
Key Responsibilities During an Incident
- The NOC Engineer should be able to follow a structured process:
- Alert → Acknowledge → Validate → Troubleshoot → Identify Impact → Escalate → Monitor Resolution → Confirm Recovery → Document/RCA
- The candidate must be comfortable working under pressure during production outages and high-priority incidents.