Talkwalker: Most-used fields#
The table below gives information about most-used fields that you can import from Talkwalker. Other fields might also be available in Adverity.
The fields that you can fetch in Adverity are updated regularly to reflect updates to data source APIs.
DEPRECATED_spam_level Metric
(DEPRECATED) An integer representing the spam level of the document, on a scale of 0 to 100.
article_extended_attributes.bluesky_shares Metric
The number of times the article has been shared on the Bluesky social media platform.
article_extended_attributes.facebook_likes Metric
The number of likes the article received on Facebook.
article_extended_attributes.facebook_reactions_total Metric
The total number of reactions (likes, loves, wows, etc.) the article received on Facebook.
article_extended_attributes.facebook_shares Metric
The number of shares the article received on Facebook.
article_extended_attributes.num_comments Metric
The total number of comments associated with the article or post.
article_extended_attributes.twitter_shares Metric
The number of shares (retweets) the article received on Twitter.
cluster_id Dimension
A unique identifier for a cluster of similar documents.
content Dimension
The full textual content of the document.
content_snippet Dimension
A snippet of the document's content, often highlighting parts that match the search query.
domain_url Dimension
The domain URL of the document's source (e.g., example.com), used for filtering across an entire domain.
engagement Metric
Measures user interaction with content, such as likes, comments, or shares.
entity_url Dimension
A list of URLs for entities (e.g., persons, brands) extracted or linked from the article.
estimated_reach Metric
The estimated number of unique users who have potentially seen the content.
extra_author_attributes.gender Dimension
The gender of the author.
extra_author_attributes.id Dimension
A unique identifier for the author of the document.
extra_author_attributes.name Dimension
The name of the author of the document.
extra_source_attributes.description Dimension
The description of the source (e.g., website, publication) of the document.
extra_source_attributes.id Dimension
A unique identifier for the content source (e.g., a specific social media account, website, or publication).
extra_source_attributes.name Dimension
The human-readable name of the content source.
extra_source_attributes.world_data.city Dimension
The city where the source of the document is located.
extra_source_attributes.world_data.continent Dimension
The continent where the source of the document is located.
extra_source_attributes.world_data.country Dimension
The country where the source of the document is located.
extra_source_attributes.world_data.country_code Dimension
The ISO 3166-1 alpha-2 country code associated with the geographic location of the content source.
extra_source_attributes.world_data.latitude Metric
The geographic latitude coordinate of the content source.
extra_source_attributes.world_data.longitude Metric
The geographic longitude coordinate of the content source.
extra_source_attributes.world_data.region Dimension
The region within the country where the source of the document is located.
extra_source_attributes.world_data.resolution Dimension
Indicates the precision or granularity of the geographical data provided for the content source (e.g., country, region, city).
fluency_level Metric
(DEPRECATED) An integer representing the fluency level of the document's language, on a scale of 0 to 100.
host_url Dimension
The host URL of the document's source (e.g., www.example.com or blog.example.com), used for filtering on specific hosts.
iab_category.0.tier1 Dimension
The primary top-level category of the content as defined by the Interactive Advertising Bureau (IAB) taxonomy.
iab_category.1.tier1 Dimension
An additional top-level category of the content, following the IAB taxonomy, for content that may fit multiple classifications.
iab_category.1.tier2 Dimension
The secondary sub-category within an additional top-level IAB category, providing a more granular classification.
images Dimension
A list of image objects associated with the document, including their URLs, legends, width, and height.
indexed Dimension
The timestamp (in milliseconds since epoch) when the document was initially indexed in the Talkwalker system.
lang Dimension
The 2-character ISO code for the language of the document's content.
noise_category Dimension
The category of noise detected in the document, such as "promotions", "hate_speech", or "job_offers".
noise_level Metric
An integer indicating the detected noise level of the document, on a scale of 0 to 100.
parent_url Dimension
The URL of the parent document in a conversation thread (e.g., the original post for a comment or retweet).
porn_level Metric
An integer indicating the detected pornographic content level of the document, on a scale of 0 to 100.
post_type Dimension
The type of post, such as "TEXT", "VIDEO", "LINK", or "AUDIO".
published Dimension
The timestamp (in milliseconds since epoch) indicating when the document was published. This field can be used for time-based filtering and histogram breakdowns.
reach Metric
The estimated number of unique people who were exposed to or reached by the article or post.
report_date Dimension
The date when the data or report was generated.
root_url Dimension
The root URL of the document's source (e.g., https://www.example.com/).
search_indexed Dimension
The timestamp (in milliseconds since epoch) indicating when the document was indexed by Talkwalker. This field can be used for time-based filtering and histogram breakdowns.
sentiment Dimension
The sentiment score of the document, determined by Talkwalker's Natural Language Processing (NLP), indicating whether the content is positive, negative, or neutral.
source_extended_attributes.alexa_pageviews Metric
The number of page views according to Alexa for the source of the document.
source_extended_attributes.alexa_unique_visitors Metric
The number of unique visitors according to Alexa for the source of the document.
source_type Dimension
The media type code of the document's source (e.g., "ONLINENEWS_NEWSPAPER", "SOCIALMEDIA").
tags_internal Dimension
Internal tags applied to the document (e.g., "hasImage", "isQuestion").
title Dimension
The title of the document.
title_snippet Dimension
A snippet of the document's title, often used for highlighting matching keywords.
url Dimension
The unique URL of the document. This field serves as the primary key.
word_count Metric
The total number of words in the document's content.