Course Outline
Day 01
Overview of Big Data Business Intelligence for Criminal Intelligence Analysis
- Law Enforcement Case Studies - Predictive Policing
- Big Data adoption rates in Law Enforcement Agencies and alignment of future operations around Big Data Predictive Analytics
- Emerging technology solutions including gunshot sensors, surveillance video, and social media
- Leveraging Big Data technology to mitigate information overload
- Integration of Big Data with Legacy data
- Foundational understanding of enabling technologies in predictive analytics
- Data Integration & Dashboard visualization
- Fraud management
- Business Rules and Fraud detection
- Threat detection and profiling
- Cost-benefit analysis for Big Data implementation
Introduction to Big Data
- Key characteristics of Big Data: Volume, Variety, Velocity, and Veracity.
- MPP (Massively Parallel Processing) architecture
- Data Warehouses – static schema, slowly evolving datasets
- MPP Databases: Greenplum, Exadata, Teradata, Netezza, Vertica, etc.
- Hadoop-Based Solutions – no structural constraints on datasets.
- Typical pattern: HDFS, MapReduce (crunch), retrieve from HDFS
- Apache Spark for stream processing
- Batch processing – suitable for analytical/non-interactive tasks
- Volume: CEP streaming data
- Typical choices – CEP products (e.g., Infostreams, Apama, MarkLogic, etc.)
- Less production-ready – Storm/S4
- NoSQL Databases – (columnar and key-value): Ideal as an analytical adjunct to data warehouses/databases
NoSQL Solutions
- KV Store - Keyspace, Flare, SchemaFree, RAMCloud, Oracle NoSQL Database (OnDB)
- KV Store - Dynamo, Voldemort, Dynomite, SubRecord, Mo8onDb, DovetailDB
- KV Store (Hierarchical) - GT.m, Cache
- KV Store (Ordered) - TokyoTyrant, Lightcloud, NMDB, Luxio, MemcacheDB, Actord
- KV Cache - Memcached, Repcached, Coherence, Infinispan, EXtremeScale, JBossCache, Velocity, Terracoqua
- Tuple Store - Gigaspaces, Coord, Apache River
- Object Database - ZopeDB, DB40, Shoal
- Document Store - CouchDB, Cloudant, Couchbase, MongoDB, Jackrabbit, XML-Databases, ThruDB, CloudKit, Prsevere, Riak-Basho, Scalaris
- Wide Columnar Store - BigTable, HBase, Apache Cassandra, Hypertable, KAI, OpenNeptune, Qbase, KDI
Data Varieties: Introduction to Data Cleaning Challenges in Big Data
- RDBMS – static structure/schema, less conducive to agile, exploratory environments.
- NoSQL – semi-structured, sufficient structure to store data without exact pre-defined schemas
- Data cleaning challenges
Hadoop
- Criteria for selecting Hadoop?
- STRUCTURED - Enterprise data warehouses/databases can store massive data (at a cost) but impose structure (limiting active exploration)
- SEMI-STRUCTURED data – challenging to manage using traditional solutions (DW/DB)
- Warehousing data = significant effort and static nature even post-implementation
- For high variety & volume of data, processed on commodity hardware – HADOOP
- Commodity hardware required to establish a Hadoop Cluster
Introduction to MapReduce /HDFS
- MapReduce – distributing computing tasks over multiple servers
- HDFS – making data locally available for computing processes (with redundancy)
- Data – can be unstructured/schema-less (unlike RDBMS)
- Developer responsibility to interpret data
- Programming MapReduce = working with Java (pros/cons), manually loading data into HDFS
Day 02
Big Data Ecosystem -- Building Big Data ETL (Extract, Transform, Load) -- Selecting the Right Big Data Tools
- Hadoop vs. Other NoSQL solutions
- Requirements for interactive, random access to data
- Hbase (column-oriented database) on top of Hadoop
- Random access to data with restrictions (max 1 PB)
- Not ideal for ad-hoc analytics, suitable for logging, counting, time-series
- Sqoop - Import from databases to Hive or HDFS (JDBC/ODBC access)
- Flume – Stream data (e.g., log data) into HDFS
Big Data Management Systems
- Moving parts, compute node start/fail: ZooKeeper - For configuration/coordination/naming services
- Complex pipeline/workflow: Oozie – managing workflows, dependencies, chaining
- Deployment, configuration, cluster management, upgrades, etc. (sys admin): Ambari
- In Cloud: Whirr
Predictive Analytics -- Fundamental Techniques and Machine Learning-Based Business Intelligence
- Introduction to Machine Learning
- Learning classification techniques
- Bayesian Prediction -- preparing training files
- Support Vector Machines
- KNN p-Tree Algebra & vertical mining
- Neural Networks
- Big Data large variable problem -- Random Forest (RF)
- Big Data Automation problem – Multi-model ensemble RF
- Automation through Soft10-M
- Text analysis tool-Treeminer
- Agile learning
- Agent-based learning
- Distributed learning
- Introduction to Open Source Tools for predictive analytics: R, Python, Rapidminer, Mahut
Predictive Analytics Ecosystem and Applications in Criminal Intelligence Analysis
- Technology and the investigative process
- Insight analytics
- Visualization analytics
- Structured predictive analytics
- Unstructured predictive analytics
- Threat/fraudster/vendor profiling
- Recommendation Engines
- Pattern detection
- Rule/Scenario discovery – failure, fraud, optimization
- Root cause discovery
- Sentiment analysis
- CRM analytics
- Network analytics
- Text analytics for extracting insights from transcripts, witness statements, internet chatter, etc.
- Technology-assisted review
- Fraud analytics
- Real-time analytics
Day 03
Real-Time and Scalable Analytics Over Hadoop
- Reasons common analytic algorithms fail in Hadoop/HDFS
- Apache Hama - for Bulk Synchronous distributed computing
- Apache SPARK - for cluster computing and real-time analytics
- CMU Graphics Lab2- Graph-based asynchronous approach to distributed computing
- KNN p -- Algebra-based approach from Treeminer for reduced operational hardware costs
Tools for eDiscovery and Forensics
- eDiscovery over Big Data vs. Legacy data – a comparison of cost and performance
- Predictive coding and Technology Assisted Review (TAR)
- Live demo of vMiner to demonstrate how TAR accelerates discovery
- Enhanced indexing through HDFS – Velocity of data
- NLP (Natural Language Processing) – open source products and techniques
- eDiscovery in foreign languages -- technology for foreign language processing
Big Data BI for Cyber Security – Comprehensive View, Rapid Data Collection, and Threat Identification
- Understanding the basics of security analytics -- attack surface, security misconfiguration, host defenses
- Network infrastructure / Large data pipeline / Response ETL for real-time analytics
- Prescriptive vs predictive – Fixed rule-based vs auto-discovery of threat rules from metadata
Aggregating Disparate Data for Criminal Intelligence Analysis
- Utilizing IoT (Internet of Things) as sensors for data capture
- Using Satellite Imagery for Domestic Surveillance
- Applying surveillance and image data for criminal identification
- Other data collection technologies -- drones, body cameras, GPS tagging systems, and thermal imaging technology
- Combining automated data retrieval with data obtained from informants, interrogation, and research
- Forecasting criminal activity
Day 04
Fraud Prevention BI from Big Data in Fraud Analytics
- Basic classification of Fraud Analytics -- rules-based vs predictive analytics
- Supervised vs unsupervised Machine Learning for Fraud pattern detection
- Business-to-business fraud, medical claims fraud, insurance fraud, tax evasion, and money laundering
Social Media Analytics -- Intelligence Gathering and Analysis
- How criminals use Social Media to organize, recruit, and plan
- Big Data ETL API for extracting social media data
- Text, image, metadata, and video
- Sentiment analysis from social media feeds
- Contextual and non-contextual filtering of social media feeds
- Social Media Dashboards to integrate diverse social media platforms
- Automated profiling of social media profiles
- Live demonstration of each analytic capability using the Treeminer Tool
Big Data Analytics in Image Processing and Video Feeds
- Image Storage techniques in Big Data -- Storage solutions for data exceeding petabytes
- LTFS (Linear Tape File System) and LTO (Linear Tape Open)
- GPFS-LTFS (General Parallel File System - Linear Tape File System) -- layered storage solution for large image data
- Fundamentals of image analytics
- Object recognition
- Image segmentation
- Motion tracking
- 3-D image reconstruction
Biometrics, DNA, and Next-Generation Identification Programs
- Advancing beyond fingerprinting and facial recognition
- Speech recognition, keystroke analysis (analyzing user typing patterns), and CODIS (Combined DNA Index System)
- Advancing beyond DNA matching: using forensic DNA phenotyping to construct faces from DNA samples
Big Data Dashboards for Quick Accessibility of Diverse Data and Display:
- Integration of existing application platforms with Big Data Dashboards
- Big Data management
- Case Study of Big Data Dashboard: Tableau and Pentaho
- Using Big Data apps to enhance location-based services in Government
- Tracking systems and management
Day 05
Justifying Big Data BI Implementation Within an Organization:
- Defining the ROI (Return on Investment) for implementing Big Data
- Case studies demonstrating time savings for analysts in data collection and preparation – increasing productivity
- Revenue gains from reduced database licensing costs
- Revenue gains from location-based services
- Cost savings from fraud prevention
- An integrated spreadsheet approach for calculating approximate expenses vs. Revenue gains/savings from Big Data implementation.
Step-by-Step Procedure for Replacing a Legacy Data System with a Big Data System
- Big Data Migration Roadmap
- Critical information needed before architecting a Big Data system?
- Methods for calculating Volume, Velocity, Variety, and Veracity of data
- Estimating data growth
- Case studies
Review of Big Data Vendors and their products.
- Accenture
- APTEAN (Formerly CDC Software)
- Cisco Systems
- Cloudera
- Dell
- EMC
- GoodData Corporation
- Guavus
- Hitachi Data Systems
- Hortonworks
- HP
- IBM
- Informatica
- Intel
- Jaspersoft
- Microsoft
- MongoDB (Formerly 10Gen)
- MU Sigma
- Netapp
- Opera Solutions
- Oracle
- Pentaho
- Platfora
- Qliktech
- Quantum
- Rackspace
- Revolution Analytics
- Salesforce
- SAP
- SAS Institute
- Sisense
- Software AG/Terracotta
- Soft10 Automation
- Splunk
- Sqrrl
- Supermicro
- Tableau Software
- Teradata
- Think Big Analytics
- Tidemark Systems
- Treeminer
- VMware (Part of EMC)
Q/A session
Requirements
- Familiarity with law enforcement processes and data systems
- Foundational knowledge of SQL/Oracle or relational databases
- Basic understanding of statistics (spreadsheet level)
Target Audience
- Law enforcement specialists with a technical background
Testimonials (3)
basics and loved the prepared documents and exercises
Rekha Nallam - GE Medical Systems Polska Sp. z o.o.
Course - Introduction to Predictive AI
Deepthi was super attuned to my needs, she could tell when to add layers of complexity and when to hold back and take a more structured approach. Deepthi truly worked at my pace and ensured I was able to use the new functions /tools myself by first showing then letting me recreate the items myself which really helped embed the training. I could not be happier with the results of this training and with the level of expertise of Deepthi!
Deepthi - Invest Northern Ireland
Course - IBM Cognos Analytics
he was well prepared - and he is very sympathetic