Is Reffer App safe?
Reffer App fetches a server-assigned target URL and CSS-selector schema, then crawls and scrapes up to 500 pages on any site.
The extension periodically polls its own backend for a task that names an arbitrary target URL plus a set of CSS selectors. Using its all-sites host permission, it fetches that URL, extracts matching text, attributes, inline script bodies, links, and images per the server-supplied schema, follows links recursively across up to 500 pages, and uploads the collected content back to the operator's servers. The client enforces no allowlist on which domains can be targeted, so the scope of any crawl is fully controlled by the backend.
Who publishes itpalbenalcaraz - 1 other listing from the same operator, 1 of them carrying a finding
palbenalcaraz - 1 other listing from the same operator, 1 of them carrying a finding
What this publisher told the store about itself, and the other listings that told it the same thing.
Same store account
1 other listing published from this account, 1k+ users between them. 1 of them carries a finding.
AI-generated. Findings may contain errors. Those marked Verified have been manually reviewed.
Publishers can request a review.
Findings
Reffer App fetches any server-named URL and crawls it with no domain limit
Code analysis shows Reffer App's backend can assign a task naming any URL.
The extension then fetches that page and up to 500 linked pages, copying text, inline scripts and links back to reffer.ai, with no domain restriction in the client.
You log into your reffer.ai account in the extension and click Start Training.
This switches the client from idle into work mode and begins polling for tasks.
The service worker fetches a task from reffer.ai's backend and visits whatever URL that task names.
The client places no restriction on which domains the backend can send it to.
| Field | Value | Why it matters | |
|---|---|---|---|
Full page text | Q3 Regional Sales Report, Internal Distribution Only | Every visible line of text on the page, extracted with the page's own text content. | |
Inline script source | function initDashboard(){ fetchAccountData(); } | The complete body of every script tag with no src attribute, including private page logic. | |
Every link and image address | https://portal.example-target.com/reports/2024-q3.pdf | The address of every link and image on the page, which also picks the next page to visit. | |
Crawl target address | https://portal.example-target.com/dashboard | The exact page the reffer.ai backend told this device to fetch for this task. |
The fetch has no domain allowlist; the crawl recurses up to 500 pages
processGetUrl: function(url, callback,callbackError){
let method = 'GET';
let contentType = 'html';
let accept = 'text/plain, */*';
let result;
let status = '0';
let d;
let started = (new Date()).getTime()/1000;
let finished = null;
let duration = null;
try {
result = fetch(url, {
method: method,
headers: {
'Accept': accept,
'Content-Type': contentType,
'X-Requested-With': 'Tasker',
},
body: null
});
result.then(response => {
status = response.status;
const r = response.text();
r.then(data=>{
finished = (new Date()).getTime()/1000;
duration = finished - started;
try{
d = JSON.parse(data);
}catch (e){
d = data;
}
data = JSON.stringify({'data': d, 'status':status, 'started': started, 'finished': finished, 'duration': duration});
if(typeof callback == 'function'){
callback(data);
}
}).catch(e => {
console.log('fetch error 1 ['+url+'] '+status, e);
if(typeof callbackError == 'function') {
callbackError('getUrlError: '+e);
}
});
}).catch(e => {
console.log('fetch error 2 ['+url+'] '+status, e);
if(typeof callbackError == 'function') {
callbackError('getUrlError: '+e);
}
});
} catch (e) {
if(typeof callback == 'function') {
callback('Request error: '+e);
}
}
return result;
}, pushToQueue(item, schema, queue){
if(domParserCounter >= maxPageQty){
log('reached maxPageQty:'+maxPageQty,{queue:queue});
return;
}
log('pushToQueue', {'item':item, 'queue':queue});
const payload = {'url': urlHelper.getUrl(item, this.url)};
if(payload.url === null){
return;
}
schema = this.getSchema(payload.url, schema);
if(schema === null){
return;
}
for(let i in schema){
payload[i] = schema[i];
}
log('pushToQueue: +');
const domParserParams = Object.assign({}, this.parameters);
domParserParams.currentLevel = this.currentLevel+1;
queue.push(domParser(payload, domParserParams));
domParserCounter++;
syslog('Download request '+payload.url);
},- app.reffer.ai
Reffer AI's own backend. Assigns crawl tasks via /app/get-tasks and receives the scraped page data via /app/save-result.
- dashboard.reffer.ai
Reffer AI's login/authorization page, listed in externally_connectable.
- operator-chosen crawl target
The page named in the task payload at runtime. Not fixed in the client and not restricted to any domain list.
Runs the extension's own page-scraping rules against any saved HTML page, so you can see exactly what a task's extraction schema pulls off a real page.
#!/usr/bin/env node
/**
* Replays the selector logic from workers/domParser/domParser.js
* (queryOne / queryMany / setDataItemValue) against a saved HTML page,
* using the extraction schema documented in the extension's own
* workers/domParser/readme.MD ("Classic declaration" example).
*/
const fs = require('fs');
const { JSDOM } = require('jsdom');
// Same rule shapes the extension's own readme documents as a task payload.
const schema = {
metaTitle: { queryOne: 'head>title', value: 'innerText' },
metaDescription: { queryOne: 'head>meta[name=description]', value: 'attr', attrName: 'content' },
scriptFiles: { queryMany: 'script:not([src=""])', value: 'attr', attrName: 'src' },
scripts: { queryMany: 'script:not([src])', value: 'innerText' },
links: { queryMany: 'body a', value: 'attr', attrName: 'href' },
images: { queryMany: 'body img', value: 'attr', attrName: 'src' },
};
// setDataItemValue() from workers/domParser/domParser.js
function setDataItemValue(element, rule) {
if (element === null) return null;
switch (rule.value) {
case 'attr':
return element.getAttribute(rule.attrName);
case 'innerText':
default:
return element.textContent;
}
}
// queryOne() / queryMany() from workers/domParser/domParser.js
function queryOne(dom, rule) {
return setDataItemValue(dom.querySelector(rule.queryOne), rule);
}
function queryMany(dom, rule) {
return Array.from(dom.querySelectorAll(rule.queryMany)).map((el) => setDataItemValue(el, rule));
}
// processData() from workers/domParser/domParser.js
function processData(dom, data) {
const result = {};
for (const key in data) {
const rule = data[key];
result[key] = 'queryOne' in rule ? queryOne(dom, rule) : queryMany(dom, rule);
}
return result;
}
const file = process.argv[2];
if (!file) {
console.error('Usage: node reffer-domparser-replay.js <path-to-saved-page.html>');
process.exit(1);
}
const html = fs.readFileSync(file, 'utf8');
const dom = new JSDOM(html).window.document;
console.log(JSON.stringify(processData(dom, schema), null, 2));
- 1node reffer-domparser-replay.js path/to/saved-page.html
Static analysis finding. This behaviour was identified by reading the shipped extension code and has not yet been reproduced in a live run. The trigger conditions and the exact data sent are read from the code, not from an observed capture.
What it can do
Permissions this extension asks for, as declared in version 2.5.7. Asking for a permission is not a finding on its own - it is what the extension can do if it chooses to.
Read and change your data on every site you visit
<all_urls> and 2 more
Store data in your browser
storage
Act on the current tab, but only after you click the extension
activeTab
Store an unlimited amount of data in your browser
unlimitedStorage
Where it sends data
Destinations our analysis observed Reffer App contacting. Sending data somewhere is not a finding on its own - an extension that syncs your settings has to talk to its own server - but it is where your data can go, and who else it goes to.
- app.reffer.ai
Reffer App sends data to app.reffer.ai. No other extension we have analysed sends data here.