Search references for DISTRIBUTED DATA-PROCESSING. Phrases containing DISTRIBUTED DATA-PROCESSING
See searches and references containing DISTRIBUTED DATA-PROCESSING!DISTRIBUTED DATA-PROCESSING
"IBM's Distributed Processing Capabilities For Large-Scale Data Base Systems, Part 1". Computerworld. Ronald G. Ross. "IBM's Distributed Processing Capabilities
Distributed_data_processing
Discrete, discontinuous representation of information
parallel distributed data processing across many commodity computers on a high bandwidth network. In such systems, the data is distributed across multiple computers
Digital_data
Distributed data processing framework
for reliable, scalable, distributed computing. It provides a software framework for distributed storage and processing of big data using the MapReduce programming
Apache_Hadoop
System with multiple networked computers
Distributed computing is a field of computer science that studies distributed systems, defined as computer systems whose inter-communicating components
Distributed_computing
Software architecture
implement a distributed file system. The designers of distributed applications must determine the best placement of the application's programs and data in terms
Distributed Data Management Architecture
Distributed_Data_Management_Architecture
Computer programming paradigm
computer science, stream processing (also known as event stream processing, data stream processing, or distributed stream processing) is a programming paradigm
Stream_processing
American software company
Automatic Data Processing, Inc. (ADP) is an American multinational provider of cloud-based human resources management, payroll processing, and professional
ADP_(company)
Open-source data analytics cluster computing framework
expected even for bug fixes. Big data Comparison of machine learning software Distributed computing Distributed data processing List of Apache Software Foundation
Apache_Spark
Computer network with multiple nodes to store information
cloud Data store Keyspace, the DDS schema Distributed hash table Distributed cache Cyber Resilience Yaniv Pessach, Distributed Storage (Distributed Storage:
Distributed_data_store
Category of programming languages
data. A data-centric programming language includes built-in processing primitives for accessing data stored in sets, tables, lists, and other data structures
Data-centric programming language
Data-centric_programming_language
Organized collection of data in computing
including data modeling, efficient data representation and storage, query languages, security and privacy of sensitive data, and distributed computing
Database
Programming paradigm in which many processes are executed simultaneously
exists. A distributed computer (also known as a distributed memory multiprocessor) is a distributed memory computer system in which the processing elements
Parallel_computing
Communications System was one of the first distributed computing platforms. The 3790 was developed by IBM's Data Processing Division (DPD) and announced in 1974
IBM_3790
Database whose data is stored in different physical locations
database Data grid Distributed cache Distributed data store Distributed hash table Routing protocol Distributed SQL "Definition: distributed database"
Distributed_database
Unified programming model for data processing pipelines
programming model to define and execute data processing pipelines, including ETL, batch and stream (continuous) processing. Beam Pipelines are defined using
Apache_Beam
Open-source distributed stream processing
information sources and manipulations to allow batch, distributed processing of streaming data. The initial release was on 17 September 2011. A Storm
Apache_Storm
Facilities containing Google servers
incrementally on a continuous basis. Later Google revealed a distributed data processing system called "Percolator" which is said to be the basis of Caffeine
Google_data_centers
Data synchronised across multiple sites
digital data is geographically spread (distributed) across many sites, countries, or institutions. In contrast to a centralized database, a distributed ledger
Distributed_ledger
Field of study to extract knowledge from data
Data science is an interdisciplinary academic field that uses statistics, scientific computing, scientific methods, processing, scientific visualization
Data_science
Type of database system
batch processing and grid computing. In addition, OLTP is often contrasted with online event processing (OLEP), which is based on distributed event logs
Online_transaction_processing
Parallel programming model
model and an associated implementation for processing and generating big data sets with a parallel and distributed algorithm on a cluster. A MapReduce program
MapReduce
Class of parallel computing applications
additional distributed data processing capabilities which are designed to run using the Hadoop MapReduce architecture. These include HBase, a distributed column-oriented
Data-intensive_computing
The International Parallel and Distributed Processing Symposium (or IPDPS) is an annual conference for engineers and scientists to present recent findings
International Parallel and Distributed Processing Symposium
International_Parallel_and_Distributed_Processing_Symposium
Software bus for high-volume data feeds
Apache Kafka is a distributed event store and stream-processing platform. It is an open-source system developed by the Apache Software Foundation written
Apache_Kafka
Computerized control system with distributed decision-making
A distributed control system (DCS) is a control system used to control industrial processes in which control functions are distributed among multiple autonomous
Distributed_control_system
Concept in probability and statistics
statistics and finds application in many fields, such as data mining and signal processing. Statistics commonly deals with random samples. A random sample
Independent and identically distributed random variables
Independent_and_identically_distributed_random_variables
Relational database programming language
defined by the Distributed Data Management Architecture. Distributed SQL processing ala DRDA is distinctive from contemporary distributed SQL databases
SQL
Centralized storage of knowledge
historic data through ETL processes that periodically migrate data from the operational systems to the warehouse. Online analytical processing (OLAP) is
Data_warehouse
Ability to seamlessly add computer resources to a given node
data, map reduce, or distributed storage system and is often associated with the infrastructure required to run large distributed sites such as Google
Hyperscale_computing
Set of events in a distributed application or protocol
Distributed data flow (also abbreviated as distributed flow) refers to a set of events in a distributed application or protocol. Distributed data flows
Distributed_data_flow
Database transaction between two or more networks
that distributed transactions are not limited to databases. The Open Group, a vendor consortium, proposed the X/Open Distributed Transaction Processing Model
Distributed_transaction
Aspect of Unisys OS 2200 operating system
so commonly used, distributed processing protocols, APIs, and development technology. The X/Open Distributed Transaction Processing model and standards
Unisys OS 2200 distributed processing
Unisys_OS_2200_distributed_processing
Manipulation of data before it is analyzed
amount of processing time. Examples of methods used in data preprocessing include cleaning, instance selection, normalization, one-hot encoding, data transformation
Data_preprocessing
Type of data structure
In distributed computing, a conflict-free replicated data type (CRDT) is a data structure that is replicated across multiple computers in a network, with
Conflict-free replicated data type
Conflict-free_replicated_data_type
Processing mode
the processing step (data load) can be quite lengthy, especially on large data volumes. This is usually remedied by doing only incremental processing, i
Online_analytical_processing
Sharing of data between running processes in a computer system
IPC mechanism. Merging data from two processes can often incur significantly higher costs compared to processing the same data on a single thread, potentially
Inter-process_communication
Scientific visualization software
using ParaView's batch processing capabilities. ParaView was developed to analyze extremely large datasets using distributed memory computing resources
ParaView
Abstract data type in computer science
and distributed memory architectures are considered. In the case of a shared memory model, the graph representations used for parallel processing are
Graph_(abstract_data_type)
Mathematical signal manipulation by computers
Digital signal processing (DSP) is the use of digital processing, such as by computers or more specialized digital signal processors, to perform a wide
Digital_signal_processing
Facility used to house computer servers
services, AI training, and large-scale data processing. By the end of 2024, there were 1,136 operational hyperscale data centers globally, a figure that doubled
Data_center
NASA program capability
capabilities transport the data to the science operations facilities. EOSDIS comprises processing facilities and Distributed Active Archive Centers across
EOSDIS
Reference model in computer science
Reference Model of Open Distributed Processing (RM-ODP) is a reference model in computer science, which provides a co-ordinating framework for the standardization
RM-ODP
Framework and distributed processing engine
stream-processing and batch-processing framework developed by the Apache Software Foundation. The core of Apache Flink is a distributed streaming data-flow
Apache_Flink
Subfield of artificial intelligence
require large data, by distributing the problem to autonomous processing nodes (agents). To reach the objective, DAI requires: A distributed system with
Distributed artificial intelligence
Distributed_artificial_intelligence
Distributed transaction processing standard
1991 by X/Open (which later merged with The Open Group) for distributed transaction processing (DTP). The goal of XA is to guarantee atomicity in "global
X/Open_XA
Open source software suite
high-performance distributed data storage and processing. It can be broadly compared to Google's GFS and MapReduce technology. Sector is a distributed file system
Sector/Sphere
In-memory data grid
a Hazelcast grid, data is evenly distributed among the nodes of a computer cluster, allowing for horizontal scaling of processing and available storage
Hazelcast
Type of parallel processing
multiple data (SIMD) is a type of parallel computing (processing) in Flynn's taxonomy. SIMD describes computers with multiple processing elements that
Single instruction, multiple data
Single_instruction,_multiple_data
Topics referred to by the same term
disc image file format Distributed Data Processing, a 1970s term referring to one of IBM's combined offerings Distributed Data Protocol, a client-server
DDP
Software architecture model
opportunities. Online event processing (OLEP) uses asynchronous distributed event logs to process complex events and manage persistent data. OLEP allows reliably
Event-driven_architecture
Type of geographic information system
people. In terms of data, the concept has been extended to include volunteered geographical information. Distributed processing allows improvements to
Distributed_GIS
Extremely large or complex datasets
Big data primarily refers to data sets that are too large or complex to be dealt with by traditional data-processing software. Data with many entries
Big_data
Distributed query engine
re-branded to Trino) is a distributed query engine for big data using the SQL query language. Its architecture allows users to query data sources such as Hadoop
Presto_(SQL_query_engine)
Method for data management
search. The challenge is magnified when working with distributed storage and distributed processing. In an effort to scale with larger amounts of indexed
Search_engine_indexing
Industrial data processing is a branch of applied computer science that covers the area of design and programming of computerized systems which are not
Industrial_data_processing
American computer scientist (born 1966)
open-source data interchange format. MapReduce, a system for large-scale data processing applications. Google File System, is a proprietary distributed file
Sanjay_Ghemawat
Repository of data stored in a raw format
or a distributed file system such as Apache Hadoop distributed file system (HDFS). There is a gradual academic interest in the concept of data lakes
Data_lake
Mechanism to allow software to execute a remote procedure
continues its process. While the server is processing the call, the client is blocked (it waits until the server has finished processing before resuming
Remote_procedure_call
Relational database management system
a subset of the data. They include components such as distributed concurrency control, flow control, and distributed query processing. The second category
NewSQL
Standard for data transmission in a local area network
Fiber Distributed Data Interface (FDDI) is a standard for data transmission in a local area network. It uses optical fiber as its standard underlying physical
Fiber Distributed Data Interface
Fiber_Distributed_Data_Interface
management, (2) global query of independent databases, and (3) distributed data processing. The word metadatabase is an addition to the dictionary. Originally
Metadatabase
Swedish computer scientist
computer scientist and entrepreneur, specializing in distributed systems, big data and data management. He is a co-founder and CEO of Databricks and
Ali_Ghodsi
Peer-to-peer Internet platform for censorship-resistant communication
censorship-resistant, anonymous communication. It uses a decentralized distributed data store to keep and deliver information, and has a suite of free software
Hyphanet
Free and open-source multi-model NoSQL database developed by Apple
software portal Ordered key-value store Database transaction Distributed database Distributed transaction List of formerly proprietary software "Releases
FoundationDB
Data-processing architecture
ordering of the data. Lambda architecture describes a system consisting of three layers: batch processing, speed (or real-time) processing, and a serving
Lambda_architecture
Any of a set of standard configurations of Redundant Arrays of Independent Disks
then recombining them. The diagram in this section shows how the data is distributed into stripes on two disks, with A1:A2 as the first stripe, A3:A4
Standard_RAID_levels
Sharing information to ensure consistency in computing
distributed concurrency control must be used, such as a distributed lock manager. Load balancing differs from task replication, since it distributes a
Replication_(computing)
High-performance computer cluster
data-parallel processing for applications utilizing big data. The HPCC platform includes system configurations to support both parallel batch data processing
HPCC
Open source platform
distributed machine learning and predictive analytics platform developed by the company H2O.ai (previously 0xdata). The software uses a distributed architecture
H2O_(software)
Transaction that reverses the effects of a prior, committed transaction
In transaction processing and distributed computing, a compensating transaction is a transaction that reverses the effects of a previously committed transaction
Compensating_transaction
Process of linking data objects in distinct models
In computing and data management, data mapping is the process of creating data element mappings between two distinct data models. Data mapping is used
Data_mapping
Computing technique employed to achieve parallelism
microarchitecture. These processors have multiple processing cores (up to 61 as of 2015) that can execute different instructions on different data. Most parallel
Multiple instruction, multiple data
Multiple_instruction,_multiple_data
Cognitive science approach
wave blossomed in the late 1980s, following a 1987 book Parallel Distributed Processing by James L. McClelland, David E. Rumelhart, et al., which introduced
Connectionism
Type of decentralized filesystem
is designed. The difference between a distributed file system and a distributed data store is that a distributed file system allows files to be accessed
Clustered_file_system
Database management system
transaction processing, and query processing. SingleStore stores relational data, JSON data, geospatial data, key-value vector data, and time series data. It
SingleStore
Computer company
to a single mass storage disc operating system and enhanced Distributed Data Processing. Proprietary operating systems included DOS and RMS (Resource
Datapoint
Parallelization across multiple processors in parallel computing environments
Data parallelism is parallelization across multiple processors in parallel computing environments. It focuses on distributing the data across different
Data_parallelism
Italian astrophysicist
the LISA Consortium, the LISA Science Team (LST), and the LISA Distributed Data Processing Center (DDPC). Sesana received his Laurea from the Università
Alberto_Sesana
The pilot implementation had a 32×32 processing element arrangement. The ICL DAP had 64×64 single bit processing elements (PEs) with 4096 bits of storage
ICL Distributed Array Processor
ICL_Distributed_Array_Processor
Computer science transaction algorithm
commitment protocol (ACP). It is a distributed algorithm that coordinates all the processes that participate in a distributed atomic transaction on whether
Two-phase_commit_protocol
Origins and events of data
Data lineage refers to the process of tracking how data is generated, transformed, transmitted and used across systems over time. It documents data's
Data_lineage
Decentralized machine learning
federated learning and distributed learning lies in the assumptions made on the properties of the local datasets, as distributed learning originally aims
Federated_learning
OLAP database developed by Kinetica DB, Inc
Kinetica is a distributed, memory-first OLAP database developed by Kinetica DB, Inc. Kinetica is designed to use GPUs and modern vector processors to improve
Kinetica_(software)
HTTP extension for collaborative editing
functionality into WebDAV, optimize processing, and eliminate the need for special-case processing. [MS-WDV]: Web Distributed Authoring and Versioning (WebDAV)
WebDAV
Memory used temporarily in data transfers
used. In a distributed computing environment, data buffers are often implemented in the form of burst buffers, which provides distributed buffering services
Data_buffer
Name used for several lines of minicomputers
Programmed Data Processor (PDP), referred to by some customers, media and authors as "Programmable Data Processor," is a term used by the Digital Equipment
Programmed_Data_Processor
Use of a GPU for computations typically assigned to CPUs
General-purpose computing on graphics processing units (GPGPU, or less often GPGP) is the use of a graphics processing unit (GPU), which typically handles
General-purpose computing on graphics processing units
General-purpose_computing_on_graphics_processing_units
Set of computers configured in a distributed computing system
Advantages include enabling data recovery in the event of a disaster and providing parallel data processing and high processing capacity. In terms of scalability
Computer_cluster
Message-passing system for parallel computers
standard for communication among processes that model a parallel program running on a distributed memory system. Actual distributed memory supercomputers such
Message_Passing_Interface
Open source distributed network software
network mesh or distributed network cluster structure with several relations types between nodes, formalize the data flow processing goes from upper node
Hierarchical Cluster Engine Project
Hierarchical_Cluster_Engine_Project
American computer scientist and software engineer
for processing and generating large datasets that became foundational for applications. MapReduce abstracts away the complexities of distributed computing
Jeff_Dean
Chinese-American businessman and computer engineer
used in data processing mode and word processing mode. They were user-programmable in data-processing mode and used the same word processing software
An_Wang
Multiprocessing memory architecture
programming distributed memory systems is how to distribute the data over the memories. Depending on the problem solved, the data can be distributed statically
Distributed_memory
Computing technique used to achieve parallelism
architectures. On distributed memory computer architectures, SPMD implementations usually employ message passing programming. A distributed memory computer
Single_program,_multiple_data
Scientific journal
sharing, and analytics; big data technologies; data visualization; architectures for massively parallel processing; data mining tools and techniques;
Journal_of_Big_Data
Classification of computer architectures
different data. MIMD architectures include multi-core superscalar processors, and distributed systems, using either one shared memory space or a distributed memory
Flynn's_taxonomy
Database management system
software Database transaction Data analysis Distributed database Distributed SQL Distributed transaction Document-oriented database Graph database Relational
Multi-model_database
Data analysis software
is an open-source framework designed for processing and analyzing large-scale spatial data in a distributed computing environment. It originated as GeoSpark
Apache_Sedona
Data structure that can be used by multiple threads
tightly coupled or a distributed collection of storage modules. Concurrent data structures, intended for use in parallel or distributed computing environments
Concurrent_data_structure
Problem easily dividable into parallel tasks
Monte Carlo method Distributed relational database queries using distributed set processing. Numerical integration Bulk processing of unrelated files
Embarrassingly_parallel
DISTRIBUTED DATA-PROCESSING
DISTRIBUTED DATA-PROCESSING
DISTRIBUTED DATA-PROCESSING
DISTRIBUTED DATA-PROCESSING
DISTRIBUTED DATA-PROCESSING
DISTRIBUTED DATA-PROCESSING
DISTRIBUTED DATA-PROCESSING
DISTRIBUTED DATA-PROCESSING
DISTRIBUTED DATA-PROCESSING