Automate Open-Source Threat Intelligence Feeds with Python & n8n

Learn to integrate and automate open-source threat intelligence feeds using a Python automation script and n8n workflows for proactive security monitoring.

Automate Open-Source Threat Intelligence Feeds with Python & n8n - Technology Tutorials
Short answer: To automate open-source threat intelligence feeds, you can use Python to interact with various OSINT APIs, fetch data, and perform initial processing. n8n then orchestrates these Python scripts, handles data storage, applies business logic, and triggers alerts or further actions, creating a reliable, low-code security automation workflow.

In the dynamic landscape of cybersecurity, staying informed about emerging threats is critical. Manually sifting through various open-source intelligence (OSINT) feeds is time-consuming and prone to human error. Developers, DevOps engineers, and small-business operators need efficient ways to gather, process, and act upon threat intelligence.

This guide provides a practical, step-by-step approach to automate OSINT feeds using Python for data acquisition and n8n for workflow orchestration. You will learn how to build an automated system that collects threat data, processes it, and integrates it into your existing security operations, enhancing your proactive defense capabilities.

What You'll Learn

  • How to identify and select relevant open-source threat intelligence feeds.
  • Developing Python scripts to interact with OSINT APIs and fetch data.
  • Integrating Python scripts into n8n workflows for automation.
  • Designing n8n workflows for data processing, filtering, and enrichment.
  • Configuring n8n to send alerts and integrate with other security tools.
  • Best practices for maintaining and scaling your automated threat intelligence system.
  • Identifying and Integrating OSINT Feeds with Python
  • Orchestrating with n8n: Building Your First Workflow
  • Alerting and Integration with Security Tools
  • Maintaining and Scaling Your Automation
  • Troubleshooting Common Issues
  • Understanding OSINT and its Value

    Open-Source Intelligence (OSINT) refers to data collected from publicly available sources to be used in an intelligence context. In cybersecurity, OSINT includes threat feeds, vulnerability databases, malware analysis reports, and dark web monitoring. These feeds provide indicators of compromise (IOCs) such as malicious IP addresses, domain names, file hashes, and URLs.

    Automating the collection and analysis of OSINT feeds offers several benefits:

    • Proactive Defense: Identify potential threats before they impact your systems.
    • Reduced Manual Effort: Free up security analysts from repetitive data collection tasks.
    • Improved Incident Response: Faster identification and correlation of IOCs during an incident.
    • Enhanced Threat Awareness: Gain a broader understanding of the threat landscape relevant to your organization.

    Why Python and n8n for Automation?

    The combination of Python and n8n provides a powerful and flexible solution for automating OSINT feeds. This approach allows you to automate OSINT feeds with Python and n8n efficiently.

    Python:

    • API Integration: Python excels at interacting with web APIs, making it ideal for fetching data from various OSINT sources.
    • Data Processing: Its extensive libraries (e.g., requests, json, pandas) are well-suited for parsing, filtering, and transforming raw threat intelligence data.
    • Flexibility: Python allows for complex logic and custom integrations that might be difficult to achieve with low-code tools alone.

    n8n:

    • Workflow Orchestration: n8n provides a visual interface to build and manage complex workflows, connecting different services and scripts.
    • Low-Code Development: It reduces the need for extensive coding, allowing for faster development and easier maintenance of automation flows.
    • Extensive Integrations: n8n has hundreds of built-in integrations for databases, messaging apps, security tools, and more, simplifying data routing and alerting.
    • Scheduling and Monitoring: Easily schedule workflows to run at specific intervals and monitor their execution.

    Setting Up Your Environment

    Before building the automation, you need to set up your development environment. This involves installing Python and its necessary libraries, and deploying n8n.

    Installing Python and Dependencies

    Ensure you have Python 3.8 or newer installed. You can download it from the official Python website. After installation, create a virtual environment for your project to manage dependencies.

    python3 -m venv osint_env
    source osint_env/bin/activate  # On Windows: osint_env\Scripts\activate
    pip install requests python-dotenv

    The requests library handles HTTP requests to APIs, and python-dotenv helps manage environment variables for API keys.

    Deploying n8n

    n8n can be deployed in several ways, including Docker, npm, or through cloud providers. For production environments, Docker is a common choice, offering portability and ease of management. Refer to the n8n installation documentation for detailed instructions.

    A typical Docker deployment might involve a docker-compose.yml file:

    version: '3.8'
    services:
      n8n:
        image: n8n.io/n8n
        restart: always
        ports:
          - "5678:5678"
        environment:
          - N8N_HOST=${N8N_HOST}
          - N8N_PORT=5678
          - N8N_PROTOCOL=http
          - WEBHOOK_URL=${WEBHOOK_URL}
          - GENERIC_TIMEZONE=America/New_York
          - TZ=America/New_York
          # For persistent data
          - N8N_USER_FOLDER=/home/node/.n8n
        volumes:
          - ~/.n8n:/home/node/.n8n

    After saving the file, run docker-compose up -d to start n8n. Access the n8n interface in your browser at http://localhost:5678 (or your configured host and port).

    Pro Tip: When deploying n8n in a production environment, ensure you configure HTTPS and strong authentication. Use a reverse proxy like Nginx or Caddy to handle SSL termination and secure access to your n8n instance.

    Identifying and Integrating OSINT Feeds with Python

    The core of your automation relies on fetching data from various OSINT sources. Python scripts will be responsible for interacting with these APIs.

    Choosing OSINT Sources

    Many OSINT sources offer free or freemium APIs. Here are a few examples:

    For this guide, we will focus on fetching IP reputation from AbuseIPDB and malware hashes from ThreatFox as examples.

    Python Script for Fetching IP Reputation Data

    Create a Python script named fetch_ip_reputation.py. This script will take an IP address as input and return its reputation from AbuseIPDB.

    import requests
    import os
    from dotenv import load_dotenv
    
    load_dotenv() # Load environment variables from .env file
    
    ABUSEIPDB_API_KEY = os.getenv("ABUSEIPDB_API_KEY")
    ABUSEIPDB_URL = "https://api.abuseipdb.com/api/v2/check"
    
    def get_ip_reputation(ip_address):
        if not ABUSEIPDB_API_KEY:
            print("ABUSEIPDB_API_KEY not set in environment variables.")
            return None
    
        params = {
            'ipAddress': ip_address,
            'maxAgeInDays': '90',
            'verbose': ''
        }
        headers = {
            'Accept': 'application/json',
            'Key': ABUSEIPDB_API_KEY
        }
    
        try:
            response = requests.get(ABUSEIPDB_URL, headers=headers, params=params)
            response.raise_for_status()  # Raise an exception for HTTP errors
            data = response.json()
            return data.get('data')
        except requests.exceptions.RequestException as e:
            print(f"Error fetching IP reputation for {ip_address}: {e}")
            return None
    
    if __name__ == "__main__":
        # Example usage (for testing outside n8n)
        test_ip = "1.1.1.1" # Example IP, replace with a known malicious IP for testing
        reputation_data = get_ip_reputation(test_ip)
        if reputation_data:
            print(f"Reputation for {test_ip}:")
            print(f"  Abuse Confidence Score: {reputation_data.get('abuseConfidenceScore')}")
            print(f"  Is Public: {reputation_data.get('isPublic')}")
            print(f"  Country Code: {reputation_data.get('countryCode')}")
            print(f"  Total Reports: {reputation_data.get('totalReports')}")
        else:
            print(f"Could not retrieve reputation for {test_ip}.")

    Create a .env file in the same directory as your Python script:

    ABUSEIPDB_API_KEY=<YOUR_ABUSEIPDB_API_KEY>

    Replace <YOUR_ABUSEIPDB_API_KEY> with your actual API key from AbuseIPDB. Remember to secure your API keys and avoid hardcoding them directly in the script.

    Python Script for Fetching Malware Hash Data

    Next, create fetch_threatfox_hashes.py to fetch recent malware hashes from ThreatFox.

    import requests
    import json
    import os
    
    THREATFOX_URL = "https://threatfox-api.abuse.ch/api/v1/"
    
    def get_recent_malware_hashes(limit=100):
        query = {
            "query": "get_iocs",
            "days": 7 # Get IOCs from the last 7 days
        }
        headers = {
            'Content-Type': 'application/json'
        }
    
        try:
            response = requests.post(THREATFOX_URL, headers=headers, data=json.dumps(query))
            response.raise_for_status()
            data = response.json()
            
            iocs = data.get('data', [])
            
            # Filter for file hashes (MD5, SHA1, SHA256)
            hash_iocs = [
                ioc for ioc in iocs 
                if ioc.get('ioc_type') in ['md5_hash', 'sha1_hash', 'sha256_hash']
            ]
            
            # Return a limited number of hashes
            return hash_iocs[:limit]
        except requests.exceptions.RequestException as e:
            print(f"Error fetching malware hashes from ThreatFox: {e}")
            return None
    
    if __name__ == "__main__":
        # Example usage
        recent_hashes = get_recent_malware_hashes(limit=5)
        if recent_hashes:
            print("Recent Malware Hashes (ThreatFox):")
            for ioc in recent_hashes:
                print(f"  Value: {ioc.get('ioc_value')}, Type: {ioc.get('ioc_type')}, First Seen: {ioc.get('first_seen_utc')}")
        else:
            print("Could not retrieve recent malware hashes.")

    This script retrieves IOCs from ThreatFox and filters them to specifically return file hashes.

    Orchestrating with n8n: Building Your First Workflow

    Now, let's integrate these Python scripts into an n8n workflow. The goal is to automate OSINT feeds with Python and n8n, process the data, and trigger actions.

    Triggering the Workflow

    In n8n, workflows begin with a trigger. For automated OSINT collection, a Cron trigger is suitable for scheduled execution.

    1. Open your n8n instance in a web browser.
    2. Click "New Workflow" or open an existing one.
    3. Add a new node and search for "Cron".
    4. Configure the Cron node:
      • Mode: "Every X"
      • Value: For example, "1 Hour" or "4 Hours" depending on how frequently you want to fetch data.

    This sets up the workflow to run at a regular interval.

    Executing Python Scripts in n8n

    n8n's Execute Command node allows you to run shell commands, including Python scripts. For this to work, ensure your Python scripts are accessible from where n8n is running (e.g., mounted volume in Docker, or present in the n8n container if built custom).

    Option 1: Python script within the n8n container (advanced)

    If you build a custom n8n Docker image, you can include your Python scripts and dependencies directly. This simplifies execution.

    Option 2: Python script on the host system (simpler for initial setup)

    If n8n is running in Docker, you can mount a volume containing your Python scripts. For example, add this to your docker-compose.yml:

        volumes:
          - ~/.n8n:/home/node/.n8n
          - ./osint_scripts:/home/node/osint_scripts # Mount your script directory

    Place your fetch_ip_reputation.py and fetch_threatfox_hashes.py (and the .env file) inside the osint_scripts directory.

    Execute Command Node Configuration:

    1. Add an Execute Command node after the Cron trigger.
    2. Configure it to run the IP reputation script:
      • Command: python3 /home/node/osint_scripts/fetch_ip_reputation.py <IP_ADDRESS_TO_CHECK> (replace <IP_ADDRESS_TO_CHECK> with a placeholder or actual IP for testing, or pass from previous node).
      • For the initial setup, you might hardcode a test IP, or fetch a list of IPs from another source (e.g., your firewall logs, a database).
      • To fetch multiple IPs, you would typically have a previous node provide a list, and then use a Split In Batches node followed by the Execute Command node.
    3. Add another Execute Command node for the ThreatFox script:
      • Command: python3 /home/node/osint_scripts/fetch_threatfox_hashes.py

    Capturing Output: The Execute Command node captures standard output (stdout) and standard error (stderr). Your Python scripts should print JSON output to stdout for easy parsing by subsequent n8n nodes.

    Modify fetch_ip_reputation.py to print JSON:

    # ... (imports and function definition) ...
    
    if __name__ == "__main__":
        # When run via n8n, it might not receive direct arguments easily for single IP
        # For a list of IPs, n8n would typically pass them in the input data.
        # For now, let's assume we want to fetch reputation for a hardcoded list or
        # get a general feed.
        
        # For demonstrating output to n8n, we will print a JSON structure.
        # In a real scenario, you'd likely iterate over IPs received from n8n.
    
        # Example: fetch for a specific IP (or multiple, if input is processed)
        ips_to_check = ["1.1.1.1", "8.8.8.8"] # Example IPs
    
        results = []
        for ip in ips_to_check:
            reputation_data = get_ip_reputation(ip)
            if reputation_data:
                results.append({
                    "ip_address": ip,
                    "abuseConfidenceScore": reputation_data.get('abuseConfidenceScore'),
                    "isPublic": reputation_data.get('isPublic'),
                    "countryCode": reputation_data.get('countryCode'),
                    "totalReports": reputation_data.get('totalReports')
                })
        
        print(json.dumps(results)) # Print JSON to stdout for n8n

    And fetch_threatfox_hashes.py already prints JSON output.

    Processing and Filtering Data in n8n

    After executing the Python scripts, the output will be available in the subsequent n8n nodes. Use JSON and Filter nodes to refine the data.

    1. JSON Node: Connect a JSON node after each Execute Command node. Set the "Data Property" to the output of the previous node (e.g., {{ $node["Execute Command"].json["stdout"] }}). This parses the JSON output from your Python script into n8n's data structure.
    2. Filter Node (for IP Reputation): Add a Filter node after the JSON node for IP reputation.
      • Condition: Filter for IPs with a high abuse confidence score. For example, {{ $json.abuseConfidenceScore > 50 }}.
      • This ensures only relevant, high-confidence malicious IPs proceed through the workflow.
    3. Filter Node (for Malware Hashes): Add a Filter node after the JSON node for ThreatFox hashes.
      • Condition: You might filter by specific malware types or only include hashes seen within the last 24 hours. For example, {{ $json.first_seen_utc.startsWith(new Date().toISOString().slice(0, 10)) }} to get today's hashes.

    Data Storage and Enrichment

    Once you have filtered relevant IOCs, you can store them or enrich them further.

    • Database Storage: Use n8n's database nodes (e.g., Postgres, MySQL, MongoDB) to store the filtered IOCs. Create a table to hold IP addresses, hashes, their confidence scores, and timestamps. This allows for historical analysis and prevents duplicate alerts.
    • Data Enrichment: You could add another Python script or n8n node to perform further enrichment. For example, for a suspicious IP, you might query a WHOIS API to get ownership details, or a GeoIP API to get its physical location.

    Pro Tip: Implement deduplication logic before storing IOCs. Use a database lookup to check if an IOC already exists. If it does, update its "last seen" timestamp rather than creating a new entry. This prevents flooding your database and alert systems with redundant information.

    Alerting and Integration with Security Tools

    The final step is to act on the filtered threat intelligence. n8n offers various integrations for alerting and connecting with other security tools.

    Sending Alerts via Slack or Email

    For immediate notification, integrate with messaging platforms or email services.

    1. Slack Node: Add a Slack node after your Filter node (for both IP and hash alerts).
      • Authentication: Configure Slack API credentials (Bot User OAuth Token).
      • Channel: Specify the Slack channel (e.g., #threat-intel-alerts).
      • Text: Construct a clear alert message using expressions to include IOC details.
        New Malicious IP Detected!
        IP: {{ $json.ip_address }}
        Confidence Score: {{ $json.abuseConfidenceScore }}
        Country: {{ $json.countryCode }}
        Total Reports: {{ $json.totalReports }}
        Source: AbuseIPDB
                        
    2. Email Node: Alternatively, use the Email Send node to send alerts.
      • Authentication: Configure SMTP server details.
      • To: Enter recipient email addresses.
      • Subject: Threat Alert: Malicious IP Detected - {{ $json.ip_address }}
      • Body: Similar to the Slack message, include relevant IOC details.

    Integrating with a SIEM or Ticketing System

    For more reliable security operations, integrate your automated OSINT feeds with a Security Information and Event Management (SIEM) system or a ticketing system.

    • SIEM Integration (e.g., Splunk, Elastic Security):
      • Use the HTTP Request node to send filtered IOCs to your SIEM's ingestion endpoint.
      • Format the data into a JSON payload that your SIEM expects.
      • Ensure proper authentication (API keys, tokens) for the SIEM endpoint.
    • Ticketing System Integration (e.g., Jira, ServiceNow):
      • Use n8n's dedicated nodes for Jira or ServiceNow (if available), or the HTTP Request node for custom API integrations.
      • Create a new incident or ticket for each high-severity IOC.
      • Populate the ticket with all relevant details (IOC value, type, source, confidence score, timestamp) to aid analysts.

    Maintaining and Scaling Your Automation

    An automated threat intelligence system requires ongoing maintenance and consideration for scalability.

    • Monitor Workflow Health: Regularly check n8n's execution logs for failed workflows or errors in Python scripts. Set up alerts within n8n for workflow failures.
    • API Key Management: Use environment variables (as demonstrated with python-dotenv) and n8n's credential management for API keys. Rotate keys periodically.
    • Error Handling: Implement reliable error handling in your Python scripts (try-except blocks) and n8n workflows (error handling branches) to gracefully manage API rate limits, network issues, or unexpected data formats.
    • Rate Limiting: Be mindful of API rate limits for each OSINT source. Implement delays in your Python scripts or use n8n's Split In Batches and Wait nodes to manage request frequency.
    • Scalability: As your needs grow, you might need to:
      • Deploy n8n on more powerful infrastructure or in a clustered setup.
      • Optimize Python scripts for performance, especially when processing large datasets.
      • Consider a dedicated message queue (e.g., RabbitMQ, Kafka) between Python scripts and n8n for high-volume data processing.
    • Regular Review: Periodically review the OSINT sources you are using. New, more effective feeds may emerge, or existing ones might become less relevant.

    Troubleshooting Common Issues

    • Python script not running in n8n:
      • Check the file path in the Execute Command node.
      • Ensure the Python executable (python3) is in the PATH of the n8n container/host.
      • Verify file permissions of your Python scripts.
      • Check n8n's execution logs for any output from stderr.
    • API key issues:
      • Double-check API keys for typos.
      • Ensure environment variables are loaded correctly (e.g., .env file is in the correct location or variables are passed to the container).
      • Verify the API key has the necessary permissions.
    • JSON parsing errors:
      • Ensure your Python script prints valid JSON to stdout.
      • Check the output of the Execute Command node in n8n's "Output" tab.
      • Verify the "Data Property" setting in the n8n JSON node.
    • Rate limit errors:
      • Consult the API documentation for specific rate limits.
      • Implement delays in your Python scripts or use n8n's Wait node to space out requests.

    Frequently Asked Questions

    What is the primary benefit of automating OSINT feeds?

    The primary benefit is proactive threat detection and reduced manual effort. Automation allows for continuous monitoring of threats, enabling faster response times and freeing up security teams to focus on more complex analysis.

    Can I use other programming languages instead of Python with n8n?

    Yes, the n8n Execute Command node can run any command-line executable, so you can use other languages like Node.js, Go, or Ruby, provided the necessary runtime is installed where n8n is running.

    How do I handle sensitive API keys in n8n?

    n8n has a built-in Credentials feature for securely storing API keys and other sensitive information. Use this feature instead of hardcoding values in nodes or passing them directly in URLs.

    What if an OSINT API changes its structure?

    If an API changes, you will need to update your Python script to reflect the new data structure. You might also need to adjust subsequent n8n nodes (like JSON or Filter nodes) that rely on the parsed data.

    Is n8n suitable for very high-volume threat intelligence processing?

    For extremely high-volume, real-time processing, n8n might require significant resource allocation or a distributed setup. For many small to medium-sized organizations, or for batch processing of feeds, it is typically well-suited. For very high-volume, consider specialized data streaming platforms in conjunction with n8n.

    How can I ensure the data from OSINT feeds is reliable?

    Reliability can be enhanced by using multiple OSINT sources and cross-referencing indicators. Implement confidence scoring and prioritize IOCs that appear in several reputable feeds. Continuously evaluate the effectiveness of your chosen feeds.

    Official documentation