graph-studio
Adding Crawlers to the Pipeline
After creating a pipeline, the next step is to add one or more crawlers. Crawlers determine what text to process.
In the pipeline, click the Crawlers tab.

Next, click the Add Input button. Graph Studio displays the Add Component dialog box. The New tab is selected and lists all available crawlers. The Existing Components tab lists crawlers that have been previously configured for other pipelines.

To add a new crawler, select the crawler. To add an existing crawler, click the Existing Components tab and select a crawler. The list below describes each of the crawlers:
- File Based Dataset Crawler: Include this crawler to process data from a file-based linked data set (FLDS) on a file store.
- Filesystem Crawler: Include this crawler to process documents, such as email messages, PDF, XML, PowerPoint, Excel, OneNote, or Word files, and images, that are available on a file store.
- Graphmart RDF Crawler: Include this crawler to process RDF in an online graphmart or specific data layer.
- Local Volume Dataset Crawler: Include this crawler to process RDF data that is stored as a linked data set (LDS) in an Graph Studio journal.
After selecting a crawler, click OK. Graph Studio opens the Create dialog box for that crawler so that you can configure it. Click a crawler name in the list below to view the details for that component:
File Based Dataset Crawler

Title: Required field that specifies the unique name for this crawler.
Description: Optional field that provides a description of this crawler.
Backing Dataset: Required field that specifies the Graph Studio dataset to crawl.
Backing Ontology: Required field that specifies the model for the dataset.
RDF Resource Type: Required field that specifies the resource type or class of data to target with this crawler.
Link Property: Optional field that specifies any link properties to crawl. A link property is a property whose value identifies the location of a linked document. When linked properties are specified, the crawler will crawl the linked documents. For example, in the triples below, fileLocation is a link property:
urn://someUnstructuredDocument urn://someProperty "file metadata" ; urn://fileLocation "/path/to/file.pdf" .
In typical use cases, this crawler is configured to define either a Link Property or a Content Property but not both.
Content Property: Optional field that identifies any content properties to crawl. A content property is a property whose value is a string literal and you want the crawler to crawl and annotate those strings. For example, in the triples below, longDescription is a content property:
urn://someUnstructuredDocument urn://someProperty "file metadata" ;
urn://longDescription "this is some interesting, likely long, unstructured text with a lot of information, and I want it to be annotated" .Base Path Connection: Required field whose value depends on whether you specified a Link Property or a Content Property:
If a Link Property was specified, the Base Path Connection is the base path to use for resolving relative file paths in the Link Property values. For example, using the example triples:
urn://someUnstructuredDocument urn://someProperty "file metadata" ; urn://fileLocation "/path/to/file.pdf" .
The
<urn://fileLocation>value of/path/to/file.pdfcould be a relative path to a location likes3://location/bucket/path/to/file.pdfor/opt/anzoshare/data/path/to/file.pdf. Therefore, the Base Path needs to be specified to resolve any relative paths and locate the linked documents.If a Content Property was specified, the Base Path Connection is a directory on the file store where the crawler can save a copy of the Content Property strings for the Graph Studio Unstructured worker instances. Saving the content to a shared file location avoids the overhead of sending the strings to the workers over the network.
Filesystem Crawler

Title: Required field that specifies the unique name for this crawler.
Description: Optional field that provides a description of this crawler.
File Crawl Location: Required field that specifies the file system crawl location. Click the field to open the File Location dialog box:

On the left side of the screen, select the storage location for the files to crawl. On the right side of the screen, navigate to the directory that contains the files. Select a directory, and then click OK.
Crawl subfolders: Optional field that specifies whether to crawl the subdirectories under the VFS Crawl Location. To crawl the subdirectories, select the Crawl subfolders checkbox. To ignore subdirectories, clear the Crawl subfolders checkbox.
Graphmart RDF Crawler

Title: Required field that specifies the unique name for this crawler.
Description: Optional field that provides a description of this crawler.
Backing Graphmart: Optional field that specifies the graphmart to crawl. To configure the grawler to crawl at the graphmart level, select one or more graphmarts in the Backing Graphmart field and leave the Backing Layer field blank.
Backing Layer: Optional field that specifies the data layer or layers that you want the pipeline to crawl. To crawl specific layers and not an entire graphmart, make sure that you leave the Backing Graphmart field blank and select the layers to crawl in the Backing Layer field. If you specify both a Backing Graphmart and a Backing Layer, the Backing Graphmart value supersedes the Backing Layer value, resulting in the entire graphmart being crawled.
Backing Ontology: Required field that specifies the model for the Backing Graphmart or Data Layer.
RDF Resource Type: Required field that specifies the resource type or class of data to target with this crawler.
Link Property: Optional field that specifies any link properties to crawl. A link property is a property whose value identifies the location of a linked document. When linked properties are specified, the crawler will crawl the linked documents. For example, in the triples below, fileLocation is a link property:
urn://someUnstructuredDocument urn://someProperty "file metadata" ; urn://fileLocation "/path/to/file.pdf" .
In typical use cases, this crawler is configured to define either a Link Property or a Content Property but not both.
Content Property: Optional field that identifies any content properties to crawl. A content property is a property whose value is a string literal and you want the crawler to crawl and annotate those strings. For example, in the triples below, longDescription is a content property:
urn://someUnstructuredDocument urn://someProperty "file metadata" ;
urn://longDescription "this is some interesting, likely long, unstructured text with a lot of information, and I want it to be annotated" .Base Path Connection: Required field whose value depends on whether you specified a Link Property or a Content Property:
If a Link Property was specified, the Base Path Connection is the base path to use for resolving relative file paths in the Link Property values. For example, using the example triples:
urn://someUnstructuredDocument urn://someProperty "file metadata" ; urn://fileLocation "/path/to/file.pdf" .
The
<urn://fileLocation>value of/path/to/file.pdfcould be a relative path to a location likes3://location/bucket/path/to/file.pdfor/opt/anzoshare/data/path/to/file.pdf. Therefore, the Base Path needs to be specified to resolve any relative paths and locate the linked documents.If a Content Property was specified, the Base Path Connection is a directory on the file store where the crawler can save a copy of the Content Property strings for the Graph Studio Unstructured worker instances. Saving the content to a shared file location avoids the overhead of sending the strings to the workers over the network.
Local Volume Dataset Crawler

Title: Required field that specifies the unique name for this crawler.
Description: Optional field that provides a description of this crawler.
Backing Dataset: Required field that specifies the Graph Studio dataset to crawl.
Backing Ontology: Required field that specifies the model for the dataset.
RDF Resource Type: Required field that specifies the resource type or class of data to target with this crawler.
Link Property: Optional field that specifies any link properties to crawl. A link property is a property whose value identifies the location of a linked document. When linked properties are specified, the crawler will crawl the linked documents. For example, in the triples below, fileLocation is a link property:
urn://someUnstructuredDocument urn://someProperty "file metadata" ; urn://fileLocation "/path/to/file.pdf" .
In typical use cases, this crawler is configured to define either a Link Property or a Content Property but not both.
Content Property: Optional field that identifies any content properties to crawl. A content property is a property whose value is a string literal and you want the crawler to crawl and annotate those strings. For example, in the triples below, longDescription is a content property:
urn://someUnstructuredDocument urn://someProperty "file metadata" ;
urn://longDescription "this is some interesting, likely long, unstructured text with a lot of information, and I want it to be annotated" .Base Path Connection: Required field whose value depends on whether you specified a Link Property or a Content Property:
If a Link Property was specified, the Base Path Connection is the base path to use for resolving relative file paths in the Link Property values. For example, using the example triples:
urn://someUnstructuredDocument urn://someProperty "file metadata" ; urn://fileLocation "/path/to/file.pdf" .
The
<urn://fileLocation>value of/path/to/file.pdfcould be a relative path to a location likes3://location/bucket/path/to/file.pdfor/opt/anzoshare/data/path/to/file.pdf. Therefore, the Base Path needs to be specified to resolve any relative paths and locate the linked documents.If a Content Property was specified, the Base Path Connection is a directory on the file store where the crawler can save a copy of the Content Property strings for the Graph Studio Unstructured worker instances. Saving the content to a shared file location avoids the overhead of sending the strings to the workers over the network.
When you have finished configuring the crawler, click Save. Graph Studio adds the crawler to the pipeline and returns to the Crawlers screen. For example:

If you want to change the crawler configuration, click the Edit icon (
) for the crawler and modify the settings as needed. If you want to add another crawler to the pipeline, repeat the steps above.When you have finished adding crawlers, follow the instructions in Add Annotators to the Pipeline to add one or more annotators to the pipeline.
Source: https://docs.sw.siemens.com/documentation/external/PL20260212925461721/en-US/graph_studio/unstructured-pipeline-crawler.htm · retrieved 2026-08-23