ATLAS-5399 Add Apache Nutch metadata model and CrawlDb/HostDB import bridge - #748
Open
lewismc wants to merge 1 commit into
Open
ATLAS-5399 Add Apache Nutch metadata model and CrawlDb/HostDB import bridge#748lewismc wants to merge 1 commit into
lewismc wants to merge 1 commit into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes were proposed in this pull request?
Catalog Apache Nutch crawls in Atlas without creating one entity per URL (ATLAS-5399).
Adds a Nutch metadata model (
addons/models/7000-Nutch/) and a CrawlDb/HostDB import bridge (addons/nutch-bridge). The bridge creates crawl, seedlist, segment, domain, and host entities, plus per-crawl host metrics onnutch_crawl_hosts. Hosts are preferred from Nutch HostDB when{crawl}/hostdb/currentexists; otherwise they are rolled up from CrawlDb. There is nonutch_urltype.Indexing lineage is not emitted here. The model includes
nutch_index_process; the Nutch Atlas IndexWriter creates that process and a generic indexDataSet(NUTCH-3210). The two sides shareatlas.cluster.nameandcrawlId.Import CLI:
The distro
nutch-hookpackage ships the bridge and its runtime dependencies so import does not need Nutch on the classpath. Docs:docs/src/documents/Hook/HookNutch.md.Out of scope: URL/page entities, Nutch REST/JobManager types (NUTCH-3165),
indexer-kafka/ATLAS_HOOK, inject/fetch/parse process types.How was this patch tested?
mvn -pl addons/nutch-bridge test(CrawlDb MapFile partitions, HostDB SequenceFile, qualified names, host rollup, model JSON)dev-support/atlas-docker(postgres) with seedhttps://nutch.apache.org/: import created crawl, seedlist, segment,apache.org,nutch.apache.org, and crawl-host metrics; Nutch IndexingJob then created the indexDataSetandnutch_index_processnutch_index_processto indexSee attached screenshot of the lineage tab for the demo indexing process.
