Skip to main content
Research Paper Graduate 5,043 words

Reducing False Positives in Anomaly Detection Using LAMS

~26 min read 7 sections Technology · Intrusion Detection System
Abstract

This paper examines the persistent problem of high false alarm rates in network anomaly-based intrusion detection systems (IDSs) and proposes the application of Local Adaptive Multivariate Smoothing (LAMS) as a mathematical remedy. The paper reviews IDS components, types, locations, and detection techniques before formally classifying false positives into structured and unstructured categories. It then presents the LAMS method, which smooths anomaly detector output by substituting each event's score with the aggregate score of comparable previously observed events. Complexity considerations, incremental model updates, and querying procedures are detailed. Experimental evaluation against NetFlow and HTTP proxy log datasets — using AUC as the accuracy metric — demonstrates that LAMS significantly reduces unstructured false alarms and can mitigate structured ones under certain conditions.

Key Takeaways
  • Introduction: History and motivation for automated intrusion detection
  • IDS Components, Types, and Challenges: Core IDS elements, misuse vs. anomaly detection tradeoffs
  • IDS Location and Architecture: Host-based, network-based, and hybrid IDS placement
  • Detection Techniques and False Positive Classification: Signature matching, anomaly detection, structured vs. unstructured false positives
  • Proposed LAMS Method: Mathematical derivation and incremental update of LAMS
  • Experimental Evaluation: NetFlow and proxy log datasets, AUC accuracy measurement
  • Conclusion: LAMS effectiveness and limitations for false positive reduction
✍️ How to write this paper — guide, tools & examples ▾

What makes this paper effective

  • Grounds the technical proposal in a clear historical narrative, tracing IDS research from Anderson (1972) through modern hybrid systems, giving readers essential context before introducing LAMS.
  • Provides a formal mathematical framework — including the Nadaraya-Watson estimator, Gaussian kernel, and incremental update equations — that makes the proposed method verifiable and reproducible.
  • Distinguishes between structured and unstructured false positives with precision, allowing the paper to make targeted claims about what LAMS can and cannot solve.
  • Supports theoretical claims with experimental evaluation across multiple real-world datasets (NetFlow records and HTTP proxy logs), using AUC as a rigorous accuracy measure.

Key academic technique demonstrated

The paper exemplifies the proposal-and-justification structure common in applied computer science research: it identifies a specific, well-documented weakness (high false alarm rates), surveys existing approaches and their shortcomings, formally defines a mathematical solution, proves convergence properties under stated assumptions, and then validates the solution empirically. The use of numbered equations alongside algorithmic pseudocode (Algorithm 1) demonstrates how to bridge theoretical mathematics and practical implementation in a single paper.

Structure breakdown

The paper opens with a historical introduction and a survey of IDS fundamentals (components, types, locations, and detection techniques), culminating in a formal classification of false positives. Section 2 presents the LAMS method with full mathematical derivations, complexity analysis, and incremental update procedures. Section 3 describes the experimental setup — including NetFlow anomaly detection, dataset construction methods, and feature selection rationale — and reports results. A brief conclusion summarizes findings and acknowledges the method's limitations regarding structured false positives.

Essay 5,043 words

Introduction

Research into detecting intrusion into computing machines began in 1972, when James Anderson produced a report for the United States Air Force identifying the need to detect intrusion into its computer networks [1]. Initially, computer system administrators believed it was sufficient to investigate logs manually to detect breaches. However, this required that administrators be highly skilled experts, and even then it was difficult for them to keep pace with rapidly developing computer systems. This need drove the automation of Intrusion Detection Systems (IDSs). The first significant research into automating IDSs was by pioneer computer security expert Anderson [2], who proposed automating intrusion detection by separating unusual or abnormal behavior in a system's audit data. A few years later, research by Denning and Neumann led to the development of the very first live detection system based on expert rules in 1985 [3]. This work laid the foundation for modern IDSs and included algorithms, manual techniques, commercial products, and rules, all designed to constantly monitor computer systems for breaches.

As computer systems have evolved over the last thirty years, their role in everyday life has expanded enormously, particularly in developed and developing countries. They have grown in importance to individuals, organizations, and governments alike. The need to detect, prevent, and respond to computer security breaches has therefore become more critical than ever. The increased use and importance of computers in the modern world has also fueled cybercrime, driven by underground organizations and governments seeking monetary or strategic advantage [4]. Consequently, reports of computer systems being hacked or breached have multiplied. Governments and cybersecurity organizations have responded by offering IDS guidelines, tutorials, and strategies to help organizations and individuals protect themselves [5].

This has resulted in modern IDS systems with widespread information-gathering and query capabilities for both logs and alerts. In practice, however, detection has focused primarily on rule- and signature-based methods — especially at the file or network level — followed by manual log analysis by experts [6, 7]. While rule- and signature-based detection systems are effective at spotting known patterns of cyberattacks, they cannot help detect new types of attacks. Moreover, they usually carry very high computational overheads. For these reasons, researchers are increasingly seeking better breach-detection systems. Research and publications have revealed significant potential in detecting breaches more efficiently through the use of mathematical algorithms.

IDS Components, Types, and Challenges

Generally, all IDSs share three main elements [1]:

  1. Data collection — All IDSs gather at least one type of data.
  2. Data conversion to feature selection — The gathered data is converted into a feature vector, a representation of its attributes.
  3. The decision engine — A system or algorithm that analyzes the feature vector and determines whether the represented event constitutes a security breach.

The decision engine is typically configured to alert the user when a problem is detected or to respond automatically. A decision engine can be classified as a hybrid detector, an anomaly detector, or a misuse detector. Misuse IDSs compare gathered data against previously defined attack patterns and flag those that match. Therefore, new attacks exploiting unidentified vulnerabilities generally go undetected. Unfortunately, misuse IDSs constitute the majority of systems on the market. Common examples include personal computer antivirus programs such as Kaspersky and McAfee, and the network-level antivirus program Snort [1]. The strongest misuse IDSs maintain large databases of known attack patterns, but because new and previously unknown signatures are discovered by manufacturers nearly every week, constant updates are required to maintain effectiveness. This is cumbersome. Moreover, misuse IDSs typically produce fewer false positives but more false negatives — also a significant problem.

With anomaly detection, normal events or behavior are learned through observation, and any significant deviations from the normal profile are reported as potential attacks. This class of IDS therefore makes it possible to identify even previously unknown attack patterns or signatures, and most anomaly detection IDSs update in real time, enabling rapid identification of attacks [8]. However, the biggest drawback is that detection accuracy is often poor — such systems frequently generate many false alarms. Furthermore, some attacks can go undetected if they are hidden within ambient data, and the training process may have a large variance. Even worse, if attack signatures are already present in the training data, the system will learn to regard them as normal behavior and produce no detection [5].

Hybrid systems are often regarded as superior because they combine the best features of other existing systems. Some proposed hybrid systems integrate misuse detection IDSs with anomaly detection IDSs [9]. When attack signature databases are available, research tends to focus on creating such hybrid systems. Modern researchers are increasingly using supervised learning algorithms to help hybrid systems learn more effectively and detect unknown attack signatures more efficiently. Supervised learning algorithms are considered less rigid than traditional systems and are thought to offer better accuracy in feature selection.

One of the most important elements determining IDS performance is feature selection. In most applications, the number of selectable features often grows exponentially, resulting in increased computational complexity, longer feature selection times, and greater training data requirements. Poor-quality features add noise and further degrade classifier performance. To address these problems, dimension reduction methods are used to identify redundancy among related features and reduce the total number of features while preserving collected information [1]. In cases where datasets have been researched further, there is a tendency to favor multiple handcrafted features and dimension reduction methods over raw data.

When the quantity of non-attack data vastly exceeds attack data, classifiers struggle to flag potential attack data. Many studies now use hybrid systems and new feature-selection mathematical algorithms to address this imbalance. While poor training data (lacking adequate attack signatures) is a problem for misuse IDSs, poor training data (containing too much noise or non-attack data) is a problem for anomaly detection IDSs. The use of statistical methods has helped isolate outliers in anomaly detection IDSs, partially addressing these challenges in hybrid systems.

IDS Location and Architecture

Intrusion Detection Systems are usually categorized by the source of information they use and their position within the network. Because an IDS's capabilities depend on the data it can access [10], its position in the network infrastructure is critically important. IDSs can be broadly divided into two categories: host-based IDSs and network-based IDSs.

Host-based intrusion detection systems (HIDSs) are programs installed on individual computer systems and are typically used to detect intrusion on a single system. Their location on the host machine provides excellent visibility into that system. However, this proximity also means they are not isolated from the host — if an attacker gains access, they can disable or mislead the IDS. Additionally, while the data available to HIDSs is context-rich, there are added expenses: the need for host access, the need to configure distributed clients, and the need to collect and manage potentially large and critical datasets from the hosts [1].

In contrast, network-based IDSs are located on separate devices, typically upstream in the network architecture, and are designed to monitor many separate systems on the same network. Network-based IDSs are generally fully isolated from the systems they monitor, making it far less likely that an attacker can access and interfere with them. However, this isolation limits the information they can gather about the systems they monitor, making it harder to detect subtle changes or abnormal events.

Hybrid distributed intrusion detection systems combine both network and host-based data into a single system, providing greater visibility of the monitored environment. The combined data is usually fed into a single decision-making algorithm or alerting process. Some hybrid distributed systems utilize virtual resources — cloud computing systems — to monitor multiple network components simultaneously. While this comes at some cost in visibility and capability, it provides isolation in case the host is compromised, and monitoring at an upstream level gives a clearer picture of activity across the entire network.

Virtual machine monitor IDSs (VMM IDSs) are generally located externally but on the same physical machine as the monitored system [11]. Some of the intrusion detection systems discussed above have been deployed in cloud systems, generally incorporating both host-based and network-based data. While such systems rely on virtual machines, they do not necessarily use the same methods as VMM IDSs.

Hybrid intrusion detection systems can also be classified as either OS-level or program-level systems. Program-level systems monitor a single application using information such as dynamic or static control flow, invoked system calls, bytecode, source code, or other application-state information. Most program-level hybrid IDSs focus on malware detection and vulnerability detection, and by extension detect anomalies and intrusions in applications.

OS-level intrusion detection systems monitor entire networks and identify abnormal patterns at the operating system level. This often involves gathering data from file system monitoring, invoked system calls, Windows Registry data, system logs, and other sources. Many such systems make efficient use of system calls to detect abnormal behavior. While patterns of system calls relating to a single application are used in program-level IDSs, the multiple patterns for different applications are combined in OS-level IDSs to enable detection of system calls across all applications and processes. Traces of system calls are used to identify repeated patterns and detect anomalies [12].

Finally, it should be noted that side-channel intrusion detection systems, which exploit physical characteristics including timings, vibrations, electromagnetic radiation, and power consumption, are gaining popularity in cybersecurity research [13] and focus particularly on physical host-level data [14, 15]. The main advantage of side-channel detection systems is their isolation from the hosts, which hinders attackers from tampering with the IDS.

3 Sections Hidden · 1,580 words
Detection Techniques and False Positive Classification380 words
Signature matching detection identifies attacks by matching data packets against predefined attack signature samples. The matching process is typically time-consuming, providing attackers with ample time…
Proposed LAMS Method620 words
The proposed Local Adaptive Multivariate Smoothing (LAMS) method aims to substitute the anomaly detector's output for each event with the mean anomaly score of comparable previous events, where the comparability or similarity between two events is captured by the kernel function. This smooths anomaly detector output and significantly reduces the rate of…
Experimental Evaluation580 words
The effectiveness of the LAMS model has been evaluated using two different anomaly identification engines: the first uses NetFlow records and the second uses HTTP proxy logs. Both anomaly identification engines are based on an ensemble of straightforward…

Conclusion

The method discussed in this work can genuinely help reduce the number of false positives produced by anomaly detection IDSs. It entails smoothing detector output over space and time, thereby enhancing the estimate of the real anomaly score. Using mild assumptions, the paper demonstrated that unstructured false alarms arising from network traffic stochasticity can be reduced significantly. It also showed that structured false alarms may be reduced in certain scenarios, though they do not have a major effect on the remaining true positives. These findings were validated using diverse datasets drawn from real network environments.

References

[1] T. R. Glass-Vanderlan, M. D. Iannacone, M. S. Vincent and Q. Chen, "A Survey of Intrusion Detection Systems Leveraging Host Data," arXiv preprint arXiv:1805.06070, 2018.

[2] J. P. Anderson, "Computer security threat monitoring and surveillance," Anderson Co., Fort Washington, PA, 1980.

[3] D. Denning and P. G. Neumann, "Requirements and model for IDES — a real-time intrusion-detection expert system," SRI International, California, 1985.

[4] M. Alazab, S. Venkatraman, P. Watters, M. Alazab and A. Alazab, "Cybercrime: the case of obfuscated malware," in Global Security, Safety and Sustainability & e-Democracy, Heidelberg, Berlin, Springer, 2012, pp. 204–211.

[5] A. K. Sood, R. Bansal and R. J. Enbody, "Cybercrime: Dissecting the state of underground enterprise," IEEE Internet Computing, vol. 17, no. 1, pp. 60–80, 2013.

[6] R. A. Bridges, J. D. Jamieson and J. W. Reed, "Setting the threshold for high throughput detectors: A mathematical approach for ensembles of dynamic, heterogeneous, probabilistic anomaly detectors," in 2017 IEEE International Conference on Big Data, Boston, MA, USA, Dec. 2017.

[7] Q. Chen and R. A. Bridge, "Automated behavioral analysis of malware: A case study of WannaCry ransomware," in 2017 16th IEEE International Conference on Machine Learning and Applications (ICMLA), Cancun, Mexico, 2017.

[8] C. R. Harshaw, R. A. Bridges, M. D. Iannacone, J. W. Reed and J. R. Goodall, "Graphprints: towards a graph analytic method for network anomaly detection," in Proceedings of the 11th Annual Cyber and Information Security Research Conference, New York, NY, 2016.

[9] K. Veeramachaneni, I. Arnaldo, V. Korrapati, C. Bassias and K. Li, "AI2: training a big data machine to defend," in 2016 IEEE 2nd International Conference on Big Data Security on Cloud (BigDataSecurity), High Performance and Smart Computing (HPSC), and Intelligent Data and Security (IDS), New York, NY, USA, 2016.

[10] C. Modi, D. Patel, B. Borisaniya, H. Patel, A. Patel and M. Rajarajan, "A survey of intrusion detection techniques in cloud," Journal of Network and Computer Applications, vol. 36, no. 1, pp. 42–57, 2013.

[11] M. Anandapriya and B. Lakshmanan, "Anomaly based host intrusion detection system using semantic based system call patterns," in 2015 9th International Conference on Intelligent Systems and Control (ISCO), Coimbatore, India, 2015.

[12] A. Kountouras, P. Kintis, C. Lever, Y. Chen, Y. Nadji, D. Dagon, M. Antonakakis and R. Joffe, Enabling Network Security Through Active DNS Datasets, Cham: Springer International Publishing, 2016.

[13] W. Choi, H. J. Jo, S. Woo, J. Y. Chun, J. Park and D. H. Lee, "Identifying ECUs through inimitable characteristics of signals in controller area networks," IEEE Transactions on Vehicular Technology, vol. 1, no. 1, p. 99, 2018.

[14] J. M. H. Jiménez, J. A. Nichols, K. Goseva-Popstojanova, S. Prowell and R. A. Bridges, "Malware detection on general-purpose computers using power consumption monitoring: A proof of concept and case study," arXiv preprint arXiv:1705.01977, 2017.

[15] J. A. Dawson, J. T. McDonald, J. Shropshire, T. R. Andel, P. Luckett and L. Hively, "Rootkit detection through phase-space analysis of power voltage measurements," in 12th IEEE International Conference on Malicious and Unwanted Software (MALCON 2017), San Juan, Puerto Rico, 2017.

[16] M. Grill, T. Pevný and M. Rehak, "Reducing false positives of network anomaly detection by local adaptive multivariate smoothing," Journal of Computer and System Sciences, vol. 83, no. 1, pp. 43–57, 2017.

[17] A. A. Nasr, M. M. Ezz and M. Z. Abdulmaged, "An intrusion detection and prevention system based on automatic learning of traffic anomalies," International Journal of Computer Network & Information Security, vol. 8, no. 1, 2016.

[18] M. Rehák and M. Grill, "Categorisation of false positives: not all network anomalies are born equal," unpublished lecture, 2014.

[19] D. van der Steeg, R. Hofstede, A. Sperotto and A. Pras, "Real-time DDoS attack detection for Cisco IOS using NetFlow," in 2015 IFIP/IEEE International Symposium on Integrated Network Management (IM), New York, USA, 2015.

[20] L. Devroye and A. Krzy, "An equivalence theorem for L1 convergence of the kernel regression estimate," Journal of Statistical Planning and Inference, vol. 23, no. 1, pp. 71–82, 1989.

[21] R. O. Duda, P. E. Hart and D. G. Stork, Pattern Classification, New York: John Wiley & Sons, 2001.

[22] F. M. Pouzols and A. Lendasse, "Adaptive kernel smoothing regression using vector quantization," in 2011 IEEE Workshop on Evolving and Adaptive Intelligent Systems (EAIS), New York, 2011.

[23] R. Tibshirani, M. Wainwright and T. Hastie, Statistical Learning with Sparsity: The Lasso and Generalizations, Chapman and Hall/CRC, 2015.

[24] J. P. Cunningham and Z. Ghahramani, "Linear dimensionality reduction: Survey, insights, and generalizations," The Journal of Machine Learning Research, vol. 16, no. 1, pp. 2859–2900, 2015.

[25] V. Fonti and E. Belitser, "Feature selection using lasso," in VU Amsterdam Research Paper in Business Analytics, Amsterdam, 2017.

[26] J. Wang and I. C. Paschalidis, "Botnet detection based on anomaly and community detection," IEEE Transactions on Control of Network Systems, vol. 4, no. 2, pp. 392–404, 2016.

[27] E. De la Hoz, E. De La Hoz, A. Ortiz, J. Ortega and B. Prieto, "PCA filtering and probabilistic SOM for network intrusion detection," Neurocomputing, vol. 164, pp. 71–81, 2015.

[28] J. F. Colom, D. Gil, H. Mora, B. Volckaert and A. M. Jimeno, "Scheduling framework for distributed intrusion detection systems over heterogeneous network architectures," Journal of Network and Computer Applications, vol. 108, pp. 76–86, 2018.

[29] A. Jakalan, J. Gong and S. Liu, "Profiling IP hosts based on traffic behavior," in 2015 IEEE International Conference on Communication Software and Networks (ICCSN), June 2015.

[30] K. Goeschel, "Reducing false positives in intrusion detection systems using data-mining techniques utilizing support vector machines, decision trees, and naive Bayes for off-line analysis," in SoutheastCon 2016, 2016.

[31] S. Garcia, M. Grill, J. Stiborek and A. Zunino, "An empirical comparison of botnet detection methods," Computers & Security, vol. 45, pp. 100–123, 2014.

Key Concepts in This Paper
LAMS Smoothing False Positives Anomaly Detection Nadaraya-Watson Estimator NetFlow Records Feature Selection Gaussian Kernel Hybrid IDS Misuse Detection Intrusion Detection
Cite This Paper
PaperDue. (2026). Reducing False Positives in Anomaly Detection Using LAMS. PaperDue. https://www.paperdue.com/study-guide/reducing-false-positives-anomaly-detection-lams-2174037

Always verify citation format against your institution’s current style guide requirements.