Name: go-bqstreamer
Owner: Kik Interactive
Description: Stream data into Google BigQuery concurrently using InsertAll()
Created: 2015-03-25 15:48:19.0
Updated: 2018-05-21 07:10:18.0
Pushed: 2017-10-29 10:36:26.0
Homepage:
Size: 100
Language: Go
GitHub Committers
User | Most Recent Commit | # Commits |
Other Committers
User | Email | Most Recent Commit | # Commits |
README
Kik and me (@oryband) are no longer maintaining this repository.
Thanks for all the contributions. You are welcome to fork and continue development.
BigQuery Streamer
Stream insert data into BigQuery fast and concurrently,
using InsertAll()
.
Features
- Insert rows from multiple tables, datasets, and projects, and insert them
bulk. No need to manage data structures and sort rows by tables -
bqstreamer does it for you.
- Multiple background workers (i.e. goroutines) to enqueue and insert rows.
- Insert can be done in a blocking or in the background (asynchronously).
- Perform insert operations in predefined set sizes, according to BigQuery's
quota policy.
- Handle and retry BigQuery server errors.
- Backoff interval between failed insert operations.
- Error reporting.
- Production ready, and thoroughly tested. We - at Rounds (now acquired by Kik) - are using it in our data gathering workflow.
- Thorough testing and documentation for great good!
Getting Started
- Install Go, version should be at least 1.5.
- Clone this repository and download dependencies:
- Version v2:
go get gopkg.in/kikinteractive/go-bqstreamer.v2
- Version v1:
go get gopkg.in/kikinteractive/go-bqstreamer.v1
- Acquire Google OAuth2/JWT credentials, so you can authenticate with BigQuery.
How Does It Work?
There are two types of inserters you can use:
SyncWorker
, which is a single blocking (synchronous) worker.
- It enqueues rows and performs insert operations in a blocking manner.
AsyncWorkerGroup
, which employes multiple background SyncWorker
s.
- The
AsyncWorkerGroup
enqueues rows, and its background workers pull and
insert in a fan-out model.
- An insert operation is executed according to row amount or time thresholds
for each background worker.
- Errors are reported to an error channel for processing by the user.
- This provides a higher insert throughput for larger scale scenarios.
Examples
Check the GoDoc examples section.
Contribute
- Please check the issues page.
- File new bugs and ask for improvements.
- Pull requests welcome!
Test
n unit tests and check coverage.
ke test
n integration tests.
is requires an active project, dataset and pem key.
port BQSTREAMER_PROJECT=my-project
port BQSTREAMER_DATASET=my-dataset
port BQSTREAMER_TABLE=my-table
port BQSTREAMER_KEY=my-key.json
ke testintegration